| ▲ | red75prime 18 hours ago | |
Yeah, I should have said RL, not RLVR. The point is RL interacts with external world, while autoregressive pretraining is limited to the passive ingestion, and RLHF relies on a model of human preferences that has no access to truth sources besides the limited training data it was built upon. | ||