| ▲ | Unearned5161 a day ago | |
If you read the thinking you can quite literally see it say "I can't just agree with all they are saying, I should find something for a constructive response". I wager that the anti-sycophancy sections in the system prompt have gotten unbalanced with the "helpful agent" parts. I imagine that the right balance will be hard to strike well given that at the end of the day we're asking the machine to have tact, and we don't quite know how to put that into an instruction yet. "Please push back when it feels right but in other cases read the room and be less rigorous" is something that plenty of humans struggle with as it is. | ||