Remix.run Logo
artninja1988 4 hours ago

That's a crazy arc 3 score. What do people think of this? Are models actually developing fluid intelligence like what the creators claim to be measuring? Is it jus do to training for it? Is the benchmark flawed?

mcbuilder 4 hours ago | parent | next [-]

Have you played Arc 3? It seems like more of a simple optimization problem (think Sokoban) than anything approaching fluid intelligence. Whether a multi hundred billion dollar company would spend time benchmaxxing a highly publicized benchmark that claims to confer AGI is an exercise left to the reader, but I doubt Claude Plays Pokemon is suddenly going to get past Mt. Doom now.

vadansky 2 hours ago | parent [-]

> Claude Plays Pokemon is suddenly going to get past Mt. Doom now.

I miss him... But for reference he did get past Doom and got pretty far in the strength puzzle too before he cut cut off. He was looping and just brute forcing it.

modeless 4 hours ago | parent | prev | next [-]

Yes, I think it indicates real progress in fluid intelligence. Clearly these models are making huge strides in usefulness which are well correlated with their ARC-AGI scores.

I don't think this is benchmaxxing. These companies are locked in a competition to produce the best software engineer, and falling behind is an existential risk. I doubt they are wasting time benchmaxxing ARC-AGI.

wyre 4 hours ago | parent | next [-]

If they were benchmaxxing, surely they would score higher than 30% on ARC-AGI.

conradkay 3 hours ago | parent [-]

Doing a quick search it seems like the average human score is 49%?

I view benchmaxxing as more of a spectrum. Mmaybe they're doing a lot more RL in environments similar to ARC-AGI 3, not even with the purpose of scoring well on any benchmark but hoping it generalizes into better performance on real, useful tasks.

dominotw 4 hours ago | parent | prev [-]

nah they could make educated guess about arc and benchmaxx it too.

an hour ago | parent | prev | next [-]
[deleted]
password54321 3 hours ago | parent | prev | next [-]

It is pretty clear at this point that current models are good at maths and problems with verifiable rewards. And puzzles are essentially math problems. Still a long way before we can say their "fluid intelligence" is effectively applicable to the real world.

criddell 2 hours ago | parent [-]

I keep wondering why there aren't more real world tests.

Maybe hook up a bunch of the AIs to a stereo camera and a couple of microphones and give them control over actuators to so they can drive cars. Then lets race them around a somewhat complex course.

When they are good enough at driving on tracks, put them on the road. Maybe see which can drive a truck with 400 cases of Coors from Texarkana, TX to Atlanta, GA and back within 28 hours.

bonoboTP 7 minutes ago | parent [-]

https://www.anthropic.com/research/claude-plays-robotics

layer8 4 hours ago | parent | prev | next [-]

It’s still “only” at 30%, and “fluid intelligence” isn’t very well-defined. The models are getting more capable, but what that means in absolute terms is anyone’s guess, because we don’t have a thorough understanding on what exactly constitutes human intelligence.

I’d say the proof is in the pudding, that is, in real-world applications. We are still seeing important limitations in LLMs.

awestroke 4 hours ago | parent | prev [-]

Doubleplus benchmaxxed