Remix.run Logo
daveyoung 3 hours ago

Two potentials from my pov:

1. Just variance in pass@K. If you prompt any model multiple times you'll see a large variance. N=1, but I find chinese open source models have a higher variance than higher-RL'd models like fable/opus.

2. They legitimately shipped a new RL checkpoint over the 7 days, which I find hard to believe.

I am leaning towards 1.

dannyw an hour ago | parent | next [-]

Or 3, they find some bug/regression in their pipeline; maybe they didn't quant parts of a model properly, maybe their inference engine had a bug, maybe some pinned MoE expert wasn't pinned, etc...

That's very plausible to have, identify, and fix in a day; especially when you get community feedback in the wild.

zarzavat 3 hours ago | parent | prev | next [-]

3. Deployment problems unrelated to the weights causing degraded performance

swiftcoder an hour ago | parent [-]

For sure the version accessible from OpenCode had a massive timeout problem the first day or so, which seemed to heavily degrade its task completion rate

re-thc 3 hours ago | parent | prev [-]

2. There was a new checkpoint. Official.

daveyoung 3 hours ago | parent [-]

do you have reference to where it was said?