Remix.run Logo
▲ astrange an hour ago

No, most of a modern LLM's training time is spent in RLVR, which does not "acquire information from an existing source". You can RL behaviors into a randomly initialized neural network.

▲hodgehog11 an hour ago | parent [-]

This is true, but you're not going to get anywhere. The pretraining phase is necessary to immensely reduce variance in the RLVR stage. Once there, RLVR has a surprising tendency to only restrict the generated space further. This is not true of RLHF, by the way, which I find to be particularly fascinating, but I digress.