Remix.run Logo
theHocineSaad 7 hours ago

As of writing this comment, Claude Opus 5 has an intelligence score of 63, not 59 (it's not the same as Gemini 3.8 Flash).

With a score of 59, Gemini 3.8 Flash is in eighth place, falling behind even Grok 4.6, Kimi k3, and GLM 5.3.

https://imgur.com/a/BMOJBED

Squarex 7 hours ago | parent | next [-]

They are all much larger and more expensive models. Google does not have a frontier model right now, but for cheap ones, they are better than event the chinese models now.

pietz 7 hours ago | parent | next [-]

That's not being debated here. The initial reported numbers were false and this was simply pointed out. You're changing the subject.

mattlondon 7 hours ago | parent | next [-]

Opus 5 medium has the same score as 3.8 flash on artificial analysis intelligence index.

Are you implying Google or Artificial Analysis are reporting false numbers? What's your source?

asdfologist 7 hours ago | parent [-]

BTW you're comparing 3.8 flash high to opus 5 medium. 3.8 flash medium scores lower.

WarmWash 5 hours ago | parent [-]

Flash models are on the order of 1/10th the size of Opus models, so some flex in the thinking level is fair.

Topfi 2 hours ago | parent [-]

Flash is just a name with no defined or consistent meaning even within labs, let alone between them. Considering both are closed weight, there is no way to truly assess how big the size delta between the two is. Then again, who cares about size, performance and end-to-end speed+cost are what matters along with task adherence, task assessment and so on.

Model size also can not be inferred by tokens/sec for a multitude of reasons, but to showcase two examples, Opus 5 and Sonnet 5, as well as Gemini 3.1 Pro Preview and 3.1 Flash have each very comparable output speeds when using the same deployment as a basis for comparison, despite it being very likely that within their generation, the former are larger than the latter. Feel the need to mention this, as I unfortunately stumble upon so many poorly reasoned, speculative hype post trying to infer model size via utterly unreliable metrics, not based in actual data.

It’s like comments below arguing about the reasoning levels not normalized to some metric (like cost, output token amount or duration) but just the labels or high, max, medium, etc. Those mean almost nothing even when comparing models based on the same pretrain (just compare GPT-5.4 to GPT-5.2), they mean less than nothing comparing different labs releases.

duplessitous 5 hours ago | parent | prev [-]

> [...] shows an intelligence score of 59, the same as Opus 5 medium!

Nothing here is false, you are simply confused. You either didn't read what they wrote in its entirety or decided to reinterpret what they did write.

knollimar 14 minutes ago | parent [-]

"Beating opus" is the false part, no?

porphyra 3 hours ago | parent | prev [-]

Better than even the Chinese models? That's a difficult-to-quantify, extremely rapidly moving target. Just today, Qwen 3.8 Max 0902 came out with a huge improvement over the previous Qwen 3.8 Max.

kamranjon 6 hours ago | parent | prev | next [-]

They said Opus 5 medium - which does have an intelligence score of 59 (you have to select it manually from the dropdown to see it)

anthonyrstevens 6 hours ago | parent | prev [-]

That 63 score is for Max. The OP specified medium.