| ▲ | sothatsit an hour ago | |
The probability values don’t really represent confidence in modern LLMs though, especially after RLHF and RLVR. System One says they use RLCD, Reinforcement Learning for Calibrated Decisions, which presumably has accurate probabilities as an explicit optimisation goal. | ||