| ▲ | tristanj 7 hours ago |
| GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot... Performance is significantly higher than Fable 5.1 Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/ |
|
| ▲ | scrlk 7 hours ago | parent | next [-] |
| Is the ARC-AGI-3 score with their custom harness? I'm guessing that is what the footnote is for? (per https://openai.com/index/how-two-settings-tripled-our-arc-ag...) |
| |
| ▲ | tedsanders 6 hours ago | parent | next [-] | | Our responses API harness just means we're using the default settings in ChatGPT and Codex, so it should more accurately reflect real world performance. We didn’t fine-tune the harness to the eval at all. ARC is reporting our score on their official leaderboard here: https://arcprize.org/leaderboard A fair ding is that the comparison with Sol is not apples-to-apples (which we footnoted in the blog), but it's because we don’t have that data. I expect Sol would score roughly 30% with the responses API harness, so the Astra improvement is more like 30% -> 99% than 8% -> 99%. Still pretty good! (I coauthored the linked blog post) | |
| ▲ | woah 6 hours ago | parent | prev | next [-] | | Haven't people demonstrated all kinds of weak LLMs getting good ARC-AGI-3 scores with special harnesses? | | | |
| ▲ | kasperni 6 hours ago | parent | prev | next [-] | | yes it is. | |
| ▲ | enraged_camel 6 hours ago | parent | prev [-] | | Yep. Incredibly misleading. Although it is not surprising at this point. They are desperate and will do anything to undermine Anthropic's upcoming IPO. | | |
|
|
| ▲ | andxor 5 hours ago | parent | prev | next [-] |
| > Performance is significantly higher than Fable 5.1 That's not clear. Need to see independent benchmarks first. |
| |
| ▲ | forgot-my-pw 5 hours ago | parent | next [-] | | We need them pelicans on bikes. | | | |
| ▲ | andxor 5 hours ago | parent | prev | next [-] | | Artificial Analysis just published their aggregate score (61). Still below Fable 5, let alone Fable 5.1. EDIT: This is suspiciously low. Calls the relevance of existing benchmarks into question. | | |
| ▲ | timpera 5 hours ago | parent | next [-] | | I agree, Opus 5 scoring higher than Fable 5 on Artificial Analysis really makes me question the relevance of these scores. | | |
| ▲ | CamperBob2 3 hours ago | parent [-] | | There is a very simple explanation for why weaker models appear to kick sand in Fable's face: Fable cannot be benchmarked because of its batshit out-of-control refusal policy. If it actually tackled all of the problems it was assigned, it would presumably kick Opus into the weeds. |
| |
| ▲ | natsucks 3 hours ago | parent | prev [-] | | I saw this too and I'm really confused. |
| |
| ▲ | forgot-my-pw 5 hours ago | parent | prev [-] | | AA benchmark: https://artificialanalysis.ai/articles/benchmarking-gpt-6-as... TLDR: it's about the same intelligence level as Opus/Fable, but it's suppose to be 70% more token efficient than GPT 5.6 Sol. So it's currently the new leader for cost efficiency frontier. |
|
|
| ▲ | leumon 7 hours ago | parent | prev | next [-] |
| The annotation on arc-agi-3 is this:
> OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations. With this configuration gpt-5.6-sol was able to reach 38,3%. So this is misleading. |
| |
| ▲ | tedsanders 5 hours ago | parent [-] | | Just to clarify, the 38.3% is on the public set, which is easier. On the private set it’s probably more like 30ish. (This hasn’t been run by ARC, so we can only estimate at the moment.) |
|
|
| ▲ | opus5_hater 6 hours ago | parent | prev | next [-] |
| any benchmark where opus 5 achieves higher scores than fable 5 in any way is not a benchmark worth trusting. |
| |
|
| ▲ | jjice 7 hours ago | parent | prev | next [-] |
| 100% on ExploitBench seems fitting given recent events. |
|
| ▲ | boutell an hour ago | parent | prev | next [-] |
| That looks more than slight. |
|
| ▲ | malshe 7 hours ago | parent | prev [-] |
| I think we need a few writing related benchmarks. |