Remix.run Logo
wxw 3 hours ago

> Scaling post-training is all we did for GLM-5.3.

Love this opening line. And wow, great results.

> As agent capability improves, much of the difficulty in scaling post-training moves from the model to the environment.

kleiba2 2 hours ago | parent | next [-]

What actually is "scaling post-training"?

FergusArgyll 2 hours ago | parent [-]

More RLVR. Give it verifiable problems, if it doesn't find a solution move on, if it does, use that as a reward signal.

Gecko4072 2 hours ago | parent [-]

Can’t this be extended quite far? Use a cerebras-served model, use verification techniques to generate and solve millions of problems and then use that as training?

gvkhna 39 minutes ago | parent [-]

That’s the whole point, just cost and compute limitations in your way (mostly).

tjwebbnorfolk 3 hours ago | parent | prev [-]

does this suggest 5.3 is the same # of parameters as 5.2?

unrvl22 an hour ago | parent | next [-]

which is the bigger headline that people don't realize. this is 744b and its head to head with Kimi K3 (2.8T), smashes DS v4 pro (1.5T). even Opus and Sol are rumored to be 1.5T+ this is half the size!

fahrradflucht 3 hours ago | parent | prev [-]

“Today we are releasing GLM-5.3. It uses the same base model as GLM-5.2 — every gain comes from post-training.“