| ▲ | wxnx 4 hours ago | ||||||||||||||||
> They are always confident, because a confident tone ranks better in RL. This makes it sound like RL rewards a confident tone -- in general, I don't think this is true (most RL is RLVR, which typically uses binary verification of correctness). I say this because the real reason "they are always confident" is in some sense even more contrived. Training text where the speaker sounded more confident is more likely to contain a correct answer. | |||||||||||||||||
| ▲ | Forgeties79 4 hours ago | parent | next [-] | ||||||||||||||||
> This makes it sound like RL rewards a confident tone Generally it does. Especially in groups. Hell look at the state of politics right now: it’s basically about being the loudest, least compromising, most confident voice in the room. It’s not just because people will assume you’re correct, it’s because if you are confidently saying something that someone wants to be right, then they’re often just going to follow it. We are all guilty of this. If I’m turning to an LLM to diagnose something medical, I am probably frustrated or uncomfortable. Maybe I’m just scared. So this magic device just instantly spits out (allegedly) exactly what is wrong and exactly what I need to do with no hesitation. I am very liable to just take it at face value because I want an answer and it gave me one, as we have seen over and over again since ChatGPT was unleashed on the world. We don’t really need to speculate, this is already a problem. | |||||||||||||||||
| ▲ | daveguy 4 hours ago | parent | prev [-] | ||||||||||||||||
> This makes it sound like RL rewards a confident tone -- in general, I don't think this is true (most RL is RLVR, which typically uses binary verification of correctness). A binary response vs rating is not related whether it learns confident or hedged tone. Either will produce a confident tone because humans respond more positively to a confident tone, hence the conman's language. Binary or not humans reward the tone and very much bias the model. But there's an even more contrived reason the training set contributes. The vast majority of human writing is confident. When the prior is greatly biased, a random number generator biased to that prior does better. The difference with humans and machines is humans are less likely to respond if they are less confident because they understand not knowing, which is why the training set is biased. It is one of the many fundamental flaw of LLM training and confusion of LLMs with intelligence. And that will not be fixed within the LLM architecture. | |||||||||||||||||
| |||||||||||||||||