| ▲ | artninja1988 4 hours ago | ||||||||||||||||||||||
That's a crazy arc 3 score. What do people think of this? Are models actually developing fluid intelligence like what the creators claim to be measuring? Is it jus do to training for it? Is the benchmark flawed? | |||||||||||||||||||||||
| ▲ | mcbuilder 4 hours ago | parent | next [-] | ||||||||||||||||||||||
Have you played Arc 3? It seems like more of a simple optimization problem (think Sokoban) than anything approaching fluid intelligence. Whether a multi hundred billion dollar company would spend time benchmaxxing a highly publicized benchmark that claims to confer AGI is an exercise left to the reader, but I doubt Claude Plays Pokemon is suddenly going to get past Mt. Doom now. | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | modeless 4 hours ago | parent | prev | next [-] | ||||||||||||||||||||||
Yes, I think it indicates real progress in fluid intelligence. Clearly these models are making huge strides in usefulness which are well correlated with their ARC-AGI scores. I don't think this is benchmaxxing. These companies are locked in a competition to produce the best software engineer, and falling behind is an existential risk. I doubt they are wasting time benchmaxxing ARC-AGI. | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | an hour ago | parent | prev | next [-] | ||||||||||||||||||||||
| [deleted] | |||||||||||||||||||||||
| ▲ | password54321 3 hours ago | parent | prev | next [-] | ||||||||||||||||||||||
It is pretty clear at this point that current models are good at maths and problems with verifiable rewards. And puzzles are essentially math problems. Still a long way before we can say their "fluid intelligence" is effectively applicable to the real world. | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | layer8 4 hours ago | parent | prev | next [-] | ||||||||||||||||||||||
It’s still “only” at 30%, and “fluid intelligence” isn’t very well-defined. The models are getting more capable, but what that means in absolute terms is anyone’s guess, because we don’t have a thorough understanding on what exactly constitutes human intelligence. I’d say the proof is in the pudding, that is, in real-world applications. We are still seeing important limitations in LLMs. | |||||||||||||||||||||||
| ▲ | awestroke 4 hours ago | parent | prev [-] | ||||||||||||||||||||||
Doubleplus benchmaxxed | |||||||||||||||||||||||