Remix.run Logo
kleiba2 2 hours ago

What actually is "scaling post-training"?

FergusArgyll 2 hours ago | parent [-]

More RLVR. Give it verifiable problems, if it doesn't find a solution move on, if it does, use that as a reward signal.

Gecko4072 2 hours ago | parent [-]

Can’t this be extended quite far? Use a cerebras-served model, use verification techniques to generate and solve millions of problems and then use that as training?

gvkhna 39 minutes ago | parent [-]

That’s the whole point, just cost and compute limitations in your way (mostly).