Remix.run Logo
rjtc 4 hours ago

I am really confused on how it can saturate ARC-AGI but still perform poorly on aggregated benchmarks:

https://artificialanalysis.ai/models

Perhaps if it was allowed this custom harness for all benchmarks it would similarily saturate?

aniviacat 3 hours ago | parent | next [-]

This benchmark gives the same intelligence score for GPT-6 Astra (max), GPT-5.6 Sol (max), and Grok 4.6 (high)? That seems very wrong to me, unless I'm misinterpreting the visualizations.

Bjorkbat 2 hours ago | parent | prev [-]

The most straightforward answer is that despite efforts to design a benchmark that, in theory, is supposed to measure generalizable intelligence, performance on ARC-AGI-3 can't be reliably correlated to performance anywhere else. I kind of lost faith in it after o1 or o3, I can't remember which, absolutely crushed ARC-AGI-1.

And, you know, maybe also some funny business. I think it's good to be a little suspicious of a model that happens to shoot upwards in performance on a specific benchmark while also kind of keeping up with the pack on a bunch of other benchmarks.