Remix.run Logo
embedding-shape 2 hours ago

Strange that the page https://artificialanalysis.ai/agents/coding-agents doesn't even mention "Qwen" once if it's now the "best" according to one of their one index?

artemisart an hour ago | parent | next [-]

They didn't run all benchmarks. It's the best in AA agentic index (GDPval-AA v2, ³-Banking) but not coding index (DeepSWE which is missing, Terminal-Bench v2.1 they have 81% vs 90% for Sol, SWE-Atlas-QnA missing).

moritzwarhier an hour ago | parent | prev | next [-]

Does "artificial analysis" mean what it says? Dubious.

But: I've been very impressed by the larger Qwen Models, and a brief try of Kimi also impressed me.

A lingering sense of quality degradation when going deep remains.

But that's not an accusation: they seem to be hitting the compute/quality tradeoff extremely well.

And on-prem capability is simply irreplaceable.

Apart from all the innovations that were driven by the strive for this optimization: quantization, "distilling" (without obvious mad-cows-disease)... I think China was an invaluable player in this progress. Intuitively, I'd even go so far to speculate that LLaMa wouldn't exist without the competition.

amelius an hour ago | parent | prev | next [-]

According to those graphs, Grok 4.5 appears to be the most cost-effective model.

user43928 37 minutes ago | parent [-]

$0.05 per task, Intelligence Index score 52 -> GPT 5.6 Luna max

$0.36 per task, Intelligence Index score 56 -> Grok 4.5 high

$1.13 per task, Intelligence Index score 58 -> Qwen 3.8 Max

$0.81 per task, Intelligence Index score 59 -> GPT 5.6 Sol xhigh

$1.80 per task, Intelligence Index score 63 -> Opus 5 xhigh

scrlk 2 hours ago | parent | prev | next [-]

Different benchmarks:

> Artificial Analysis Agentic Index: Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, Tau³-Banking)

> Artificial Analysis Coding Agent Index v1.3 incorporates 3 benchmarks: DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA

Qwen3.8 Max is 55.4 on the Agentic Index but hasn't been tested for the Coding Agent Index.

apitman an hour ago | parent [-]

Looks like coding agent is model+harness. There are far fewer models represented on that page. I believe "agentic index" is still the metric to look at for coding performance. I could be wrong about that though.

Bootvis an hour ago | parent | prev [-]

Indeed, and this Qwen 3.8 max specific page:

https://artificialanalysis.ai/models/qwen3-8-max

Doesn't have the claim either. Clickbait?

petu an hour ago | parent [-]

This page has it, scroll to "Intelligence" header (not the highlights one, but second on the page / with black square) and click "Agentic Index"

Bootvis an hour ago | parent [-]

So the original link should be: https://artificialanalysis.ai/models/qwen3-8-max?intelligenc...

Even then, this seems a much more marginal win than the headline suggested to me.