Remix.run Logo
▲ Ariarule a day ago

Always glad to see more open-weight models, but this caption on the 2nd demo image had me do a double-take: "Land or Water Generalization Experiment: We recreated the viral X puzzle by asking Beam to create a fixed 180×90 grid for longitudes -179° to 179° and latitudes -89° to 89°, with 16,200 points. This puzzle is a few days old, so could not appear in the training data, thus testing the model’s generalization. Beam gets 95.5% coverage right, putting us between Opus 5 (92.5%) and Fable 5 (97.8%), which shows how well it generalizes to novel new tasks."

Oof, no, this "puzzle is a few days old" is incorrect even if it's a social media trend just recently. Asking a model to generate a world map in this way is _at least_ from August 2025 as it appeared on LessWrong at that time: https://www.lesswrong.com/posts/xwdRzJxyqFqgXTWbH/how-does-a...

▲extr a day ago | parent | next [-]

Yeah I remember when the original post about this came out. Def not recent. Though I think their point survives in that they didn't exactly RL on this.

▲glitchc a day ago | parent | prev | next [-]

It could be old but still not be part of the training set.

▲stingraycharles 13 hours ago | parent [-]

Yes, but then you can’t use “it’s 2 days old” to assert that it can’t be in the training set.

Seems sloppy.

▲jiggawatts a day ago | parent | prev | next [-]

I'd love to see these tests repeated for the current frontier models...

▲az226 18 hours ago | parent | prev | next [-]

Seems rather sloppy to not validate it wasn’t in their training or RL data.

▲charlieyu1 a day ago | parent | prev [-]

I don't think age of the puzzle even matters, all models have search capacities these days

▲criemen a day ago | parent | next [-]

> all models have search capacities these days

one would hope that they disable websearch and internet access (maybe all tools?) when doing generalization testing?

▲zakisaad a day ago | parent | prev | next [-]

Model weights (what is being tested here) don't inherently "access the web" when inference is running. If the model has access to a web search tool, that's a different story.

▲dexwiz a day ago | parent | prev | next [-]

Is search part of the model or the harness?

▲Cycl0ps a day ago | parent [-]

Maybe that's a rhetorical question but just in case - the search would always be part of the harness. A model is only handling next-token prediction for a given input. That token may be something like [[web search]] to invoke a tool call but the actual call would be handled by the harness.

▲dexwiz a day ago | parent | next [-]

Yeah it was rhetorical. Search would be implemented as a tool call. Pure intelligence tests would likely have limited tools. But maybe they would have a python sandbox to solve issues like Rs in strawberry.

▲ a day ago | parent | prev [-]
[deleted]
▲Barbing a day ago | parent | prev [-]

If they made that statement and knowingly had search enabled, it would essentially be fraudulent.