Remix.run Logo
woah 7 hours ago

Haven't people demonstrated all kinds of weak LLMs getting good ARC-AGI-3 scores with special harnesses?

tintor 6 hours ago | parent [-]

Those people haven't verified their results against the private set: https://arcprize.org/leaderboard

andriy_koval 6 hours ago | parent [-]

Astra also not verified using private set, but on "semi-private" set

andrewchambers 4 hours ago | parent [-]

if that is true then why is astra on the official ARC leaderboard now ?

andriy_koval 4 hours ago | parent [-]

ARC leaderboard has results from semi-private data for frontier models, they have another competition for private data.

It is described in their methodology: https://arcprize.org/policy

It makes sense, since once OpenAI API receive task, it is not private anymore but leaked to OpenAI.