| ▲ | softwaredoug 7 hours ago |
| I'm seeing reporting it gets 98.6% on ARC-AGI3[1] (previously like 30% with Fable) https://venturebeat.com/technology/welcome-to-the-agi-era-op... |
|
| ▲ | aabhay 7 hours ago | parent | next [-] |
| This is with the caveat that OpenAI uses their own harness for this: > On ARC-AGI-3, GPT-6 Astra was run with our responses API harness , which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3. |
| |
| ▲ | Readerium 6 hours ago | parent | next [-] | | Its 62 percent when using a neutral harness.
https://arcprize.org/blog/astra | | |
| ▲ | sbinnee 5 hours ago | parent | next [-] | | Yet it is an impressive number. But yeah when you see a number 99 you have doubts. Thanks for the link | |
| ▲ | glenstein 5 hours ago | parent | prev [-] | | Interesting both this and Sol got approximately a 37% boost with the custom harness. |
| |
| ▲ | simianwords 6 hours ago | parent | prev [-] | | This should be normalised and expected - the responses API harness allows it to use the custom compaction that is not allowed otherwise. It is entirely fair to allow OpenAI to use their own compaction algorithm.. | | |
| ▲ | ActionHank 6 hours ago | parent [-] | | "This should be allowed, let me explain the reason they cheated and state again that they should be allowed to cheat." | | |
| ▲ | simianwords 6 hours ago | parent [-] | | > GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game. > Going forward, we will report both Standard harness and Provider Adapter harness results on the ARC-AGI leaderboard, with each evaluation condition clearly labeled. Our open-source testing repository and testing policy document both approaches. This is what the Author of the benchmark has to stay. Quality of the comments keep going down smh | | |
| ▲ | ActionHank 6 hours ago | parent [-] | | "Going forward we will capitulate and still try to keep the integrity of our benchmark in tact, but from now on every benchmark will be compromised with providers being able tweak things sufficiently to game at least a 30% bump in results." | | |
| ▲ | simianwords 6 hours ago | parent [-] | | "I'll twist the words of the author of the benchmark itself to make a point" | | |
| ▲ | ActionHank 5 hours ago | parent [-] | | "I refuse to see the wall that I am running directly into, because if I see it I will hit it" | | |
| ▲ | simianwords 5 hours ago | parent [-] | | If you mean a capability wall, the author of the benchmark says this >We see Astra as a major breakthrough in model intelligence. You think the author of the benchmark is also in the conspiracy |
|
|
|
|
|
|
|
|
| ▲ | kasperni 7 hours ago | parent | prev | next [-] |
| "On the current ARC-AGI-3 leaderboard, conventional frontier-model runs sit dramatically below Astra's reported 98.6% result. But the comparison isn't straightforward. OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations." |
|
| ▲ | arctic-true 7 hours ago | parent | prev | next [-] |
| The blog post says 99.9%. Oddly, it does better on ARC-AGI-3 than it does on version 1 or 2 of the same benchmark (though gets 95+ on all three) |
| |
| ▲ | _diyar 7 hours ago | parent [-] | | I strongly suspect that is way above the human average anyway, esp. ARC 2 and 3 are really tough unless you happen to be great at those spacial puzzles or video games. | | |
| ▲ | aesthesia 6 hours ago | parent | next [-] | | Scoring for ARC-AGI-3 is constructed so that the median(-ish) human score is 100%, so this is not a superhuman result. However, the scaling is weird, since it's built from terms that look like (AI turns taken / median human turns) ^ 2, and it weights later levels higher than early levels. So it's not at all clear that 100% is twice as good as 50%. | | | |
| ▲ | _superposition_ 6 hours ago | parent | prev | next [-] | | Really though? I would believe something like this if a model could one shot every solution in the set. I don't pay much attention to these things and maybe this stuff is available but I would bet the session/reasoning transcript is absolutely horrendous from an intelligence standpoint. | |
| ▲ | CamperBob2 7 hours ago | parent | prev [-] | | At this point the only valid ARC-AGI benchmark left is to make up the next series of ARC-AGI benchmark puzzles that current models presumably can't handle. | | |
| ▲ | jaggederest 6 hours ago | parent [-] | | I feel like making a human-proof benchmark is pretty clear evidence that they've exceeded even the highest human capacity in most respects, for things that you can do via text generation (and to a lesser extent image generation) |
|
|
|
|
| ▲ | Bluestein 7 hours ago | parent | prev [-] |
| 100%, some say.- |