Remix.run Logo
myworkaccount2 4 hours ago

IMO the HLE scores without tools seem to align better with real world performance of the models.

To me it feels like the difference between "RL performance" and the pretraining / base "knowledge".

Yes you can RL terminal bench to the moon but does the model hold up on out of distribution tasks?

Kind of like trying to navigate a dark room with a laser light, vs a flashlight. Laser is going to go a lot farther much more efficiently but only if you are already pointing it at the right place.