| ▲ | Alpha3031 4 days ago | |
I feel like RLHF has a pretty obvious ground truth, human feedback is used as an (albeit noisy) signal of average human preferences. Same thing with RLVR and "solving the problem". | ||
| ▲ | garrinm 3 days ago | parent [-] | |
To be more specific there’s no ground truth tokens to predict. There a verifiable answer in RLVR. But the tokens are explored. Not predicted as there’s no true token to predict. | ||