Remix.run Logo
fsh a day ago

I would be very surprised if any of the frontier models wasn't trained on all public physics benchmarks. Training data providers have been hiring people for exactly this task.

letmevoteplease a day ago | parent | next [-]

This study appears to be evidence against that: the model failed the benchmark but arrived at the correct answer.

fsh a day ago | parent [-]

Half of the benchmarks don't have published answers. That's why training data providers have been hiring physicists to solve them.

red75prime a day ago | parent | next [-]

This is unconventional benchmaxxing then, when they decrease the benchmark scores to allow models to generalize on correct solutions.

a day ago | parent | prev [-]
[deleted]
bobmarleybiceps a day ago | parent | prev | next [-]

yeah, it would be almost shocking if an open source benchmark was NOT used ~somewhere in training. Perhaps just pre-training, but still. Neural networks can be fairly robust to some mistakes in their training data, so maybe it doesn't even matter if some of them are incorrect. Who knows.

redwood a day ago | parent | prev [-]

I'd have thought the same but this article from yesterday blew my mind https://www.amazon.science/blog/why-dont-machine-learning-re...

As it essential implies that these models compress knowledge well in a way that what remains is what's generalizeable more so than remembering every specific detail... Anyway more understanding necessary but thought provoking