Remix.run Logo
modeless 4 hours ago

Yes, I think it indicates real progress in fluid intelligence. Clearly these models are making huge strides in usefulness which are well correlated with their ARC-AGI scores.

I don't think this is benchmaxxing. These companies are locked in a competition to produce the best software engineer, and falling behind is an existential risk. I doubt they are wasting time benchmaxxing ARC-AGI.

wyre 4 hours ago | parent | next [-]

If they were benchmaxxing, surely they would score higher than 30% on ARC-AGI.

conradkay 3 hours ago | parent [-]

Doing a quick search it seems like the average human score is 49%?

I view benchmaxxing as more of a spectrum. Mmaybe they're doing a lot more RL in environments similar to ARC-AGI 3, not even with the purpose of scoring well on any benchmark but hoping it generalizes into better performance on real, useful tasks.

dominotw 4 hours ago | parent | prev [-]

nah they could make educated guess about arc and benchmaxx it too.