Remix.run Logo
tristanj 7 hours ago

GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot...

Performance is significantly higher than Fable 5.1

Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/

scrlk 7 hours ago | parent | next [-]

Is the ARC-AGI-3 score with their custom harness? I'm guessing that is what the footnote is for? (per https://openai.com/index/how-two-settings-tripled-our-arc-ag...)

tedsanders 6 hours ago | parent | next [-]

Our responses API harness just means we're using the default settings in ChatGPT and Codex, so it should more accurately reflect real world performance. We didn’t fine-tune the harness to the eval at all.

ARC is reporting our score on their official leaderboard here: https://arcprize.org/leaderboard

A fair ding is that the comparison with Sol is not apples-to-apples (which we footnoted in the blog), but it's because we don’t have that data. I expect Sol would score roughly 30% with the responses API harness, so the Astra improvement is more like 30% -> 99% than 8% -> 99%. Still pretty good!

(I coauthored the linked blog post)

woah 6 hours ago | parent | prev | next [-]

Haven't people demonstrated all kinds of weak LLMs getting good ARC-AGI-3 scores with special harnesses?

tintor 6 hours ago | parent [-]

Those people haven't verified their results against the private set: https://arcprize.org/leaderboard

andriy_koval 5 hours ago | parent [-]

Astra also not verified using private set, but on "semi-private" set

andrewchambers 3 hours ago | parent [-]

if that is true then why is astra on the official ARC leaderboard now ?

andriy_koval 3 hours ago | parent [-]

ARC leaderboard has results from semi-private data for frontier models, they have another competition for private data.

It is described in their methodology: https://arcprize.org/policy

It makes sense, since once OpenAI API receive task, it is not private anymore but leaked to OpenAI.

kasperni 6 hours ago | parent | prev | next [-]

yes it is.

enraged_camel 6 hours ago | parent | prev [-]

Yep. Incredibly misleading. Although it is not surprising at this point. They are desperate and will do anything to undermine Anthropic's upcoming IPO.

10xDev 6 hours ago | parent [-]

It is about memory retention. No heavy lifting done on the reasoning side so I hardly see anything misleading here.

Edit: update from fchollet https://x.com/fchollet/status/2095598451115614371

andxor 5 hours ago | parent | prev | next [-]

> Performance is significantly higher than Fable 5.1

That's not clear. Need to see independent benchmarks first.

forgot-my-pw 5 hours ago | parent | next [-]

We need them pelicans on bikes.

bwat49 5 hours ago | parent [-]

Its time to move on to the flamingo on a unicycle bench

andxor 5 hours ago | parent | prev | next [-]

Artificial Analysis just published their aggregate score (61).

Still below Fable 5, let alone Fable 5.1.

EDIT: This is suspiciously low. Calls the relevance of existing benchmarks into question.

timpera 5 hours ago | parent | next [-]

I agree, Opus 5 scoring higher than Fable 5 on Artificial Analysis really makes me question the relevance of these scores.

CamperBob2 3 hours ago | parent [-]

There is a very simple explanation for why weaker models appear to kick sand in Fable's face: Fable cannot be benchmarked because of its batshit out-of-control refusal policy.

If it actually tackled all of the problems it was assigned, it would presumably kick Opus into the weeds.

natsucks 3 hours ago | parent | prev [-]

I saw this too and I'm really confused.

forgot-my-pw 5 hours ago | parent | prev [-]

AA benchmark: https://artificialanalysis.ai/articles/benchmarking-gpt-6-as...

TLDR: it's about the same intelligence level as Opus/Fable, but it's suppose to be 70% more token efficient than GPT 5.6 Sol. So it's currently the new leader for cost efficiency frontier.

leumon 7 hours ago | parent | prev | next [-]

The annotation on arc-agi-3 is this: > OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations.

With this configuration gpt-5.6-sol was able to reach 38,3%. So this is misleading.

tedsanders 5 hours ago | parent [-]

Just to clarify, the 38.3% is on the public set, which is easier. On the private set it’s probably more like 30ish. (This hasn’t been run by ARC, so we can only estimate at the moment.)

opus5_hater 6 hours ago | parent | prev | next [-]

any benchmark where opus 5 achieves higher scores than fable 5 in any way is not a benchmark worth trusting.

machomaster 5 hours ago | parent | next [-]

Why would Anthropic trust and use these tests in their official comparisons?

ActionHank 5 hours ago | parent | prev | next [-]

username checks out

r_lee 5 hours ago | parent | prev [-]

great username lol

jjice 7 hours ago | parent | prev | next [-]

100% on ExploitBench seems fitting given recent events.

boutell an hour ago | parent | prev | next [-]

That looks more than slight.

malshe 7 hours ago | parent | prev [-]

I think we need a few writing related benchmarks.