| ▲ | StilesCrisis 5 hours ago | |||||||||||||||||||||||||||||||
LLMs are famously bad at determining "if it's not sure it is correct." They are always confident, because a confident tone ranks better in RL. | ||||||||||||||||||||||||||||||||
| ▲ | wxnx 4 hours ago | parent [-] | |||||||||||||||||||||||||||||||
> They are always confident, because a confident tone ranks better in RL. This makes it sound like RL rewards a confident tone -- in general, I don't think this is true (most RL is RLVR, which typically uses binary verification of correctness). I say this because the real reason "they are always confident" is in some sense even more contrived. Training text where the speaker sounded more confident is more likely to contain a correct answer. | ||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||