Remix.run Logo
GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index(artificialanalysis.ai)
23 points by wertyk 15 hours ago | 14 comments
bashtoni 11 hours ago | parent | next [-]

Are the Artificial Analysis benchmarks really worthwhile any more?

They really don't seem to match my real-world experience, and based on the comments I see I don't think that match most other people's either.

For example, Opus 5 was at the top for some time. My experience is that it's not noticeably better than Opus 4.8, and it definitely seems worse than Fable 5, which AA benchmarks put behind Opus 5. GPT 5.6-sol and Opus 5 seem pretty interchangeable, although Sol is noticeably better at finding problems in code, particularly edge cases.

villish 7 hours ago | parent [-]

I have no faith in these benchmarks. Muse 1.3 shouldn’t even be in the same conversation yet it scores above GPT 5.6 & 6.0

ekojs 14 hours ago | parent | prev | next [-]

Well, seems like ECI [0] and the AA index is diverging quite a bit. Benchmarking LLM is tough and I think we are seeing the limitations of current benchmarks and applicability to real tasks.

[0]: https://x.com/EpochAIResearch/status/2095602754282783108

aogaili 9 hours ago | parent | prev | next [-]

I don't understand how those labs are releasing models so close in performance to one another?

Are they just scaling more? getting more data at the same rate? training against the same benchmarks? making the same breakthroughs?

How can this be explained?

akie 4 hours ago | parent [-]

We're hitting a limit, that's what's happening.

I believe the development of new models has been like an s-curve: There were enormous incremental improvements earlier on, but now we're reaching the right-hand side of the s-curve and all the new models are clustering together.

Even the lighter and smaller and cheaper models going to end up near that limit, and that will likely evaporate the perceived economic value of both OpenAI and Anthropic unless they manage to lock it in/offset it with platform effects and branding (they might well be able to do that).

My impression is that we might well be able to move past that limit, but we would need another radical invention like the transformer architecture that the whole current generation of models is built on.

NiekvdMaas 15 hours ago | parent | prev | next [-]

Title: "major gains"

First chart: from score 61 (GPT-5.6 Sol) to drumroll 61 (GPT-6 Astra)

flyaway123 14 hours ago | parent | next [-]

Indeed. Though to be fair it is referring to "Artificial Analysis Coding Agent Index", from 65 to 67.

dist-epoch 14 hours ago | parent | prev [-]

I think they mean cost per task, where Astra is now on the Pareto frontier.

x3haloed 13 hours ago | parent | prev | next [-]

Fascinating. This is the only benchmark I've seen so far with lack-luster results. I don't understand enough about AA's specific methodology to get the implications.

usaar333 13 hours ago | parent [-]

It's sub-Fable 5 on mirror code: https://epoch.ai/benchmarks/mirrorcode?view=graph&tab=leader...

eis 14 hours ago | parent | prev | next [-]

In the general Intelligence Index it scores exactly equal to Sol (61). In the Agentic Index it scores significantly lower than Sol (51 vs 58). In both it scores lower than Fable 5.1, Opus 5 and even Muse Spark 1.3.

Am I missing something or is this not looking too... stellar?

Readerium 15 hours ago | parent | prev [-]

more like 5.7 not 6

JV00 5 hours ago | parent | next [-]

It’s based on a new and larger pre train if I understand correctly, hence the major version.

lostmsu 14 hours ago | parent | prev [-]

5.6.1