Remix.run Logo
pietz 8 hours ago

That's not being debated here. The initial reported numbers were false and this was simply pointed out. You're changing the subject.

mattlondon 8 hours ago | parent | next [-]

Opus 5 medium has the same score as 3.8 flash on artificial analysis intelligence index.

Are you implying Google or Artificial Analysis are reporting false numbers? What's your source?

asdfologist 8 hours ago | parent [-]

BTW you're comparing 3.8 flash high to opus 5 medium. 3.8 flash medium scores lower.

WarmWash 6 hours ago | parent [-]

Flash models are on the order of 1/10th the size of Opus models, so some flex in the thinking level is fair.

nomel an hour ago | parent | next [-]

When comparing closed models, the only thing that actually matters to anyone using them is some mix of cost and speed. Considering how much memory a server is using, when evaluating models that you'll never have access to, to host yourself, doesn't really make sense.

Topfi 4 hours ago | parent | prev [-]

Flash is just a name with no defined or consistent meaning even within labs, let alone between them. Considering both are closed weight, there is no way to truly assess how big the size delta between the two is. Then again, who cares about size, performance and end-to-end speed+cost are what matters along with task adherence, task assessment and so on.

Model size also can not be inferred by tokens/sec for a multitude of reasons, but to showcase two examples, Opus 5 and Sonnet 5, as well as Gemini 3.1 Pro Preview and 3.1 Flash have each very comparable output speeds when using the same deployment as a basis for comparison, despite it being very likely that within their generation, the former are larger than the latter. Feel the need to mention this, as I unfortunately stumble upon so many poorly reasoned, speculative hype post trying to infer model size via utterly unreliable metrics, not based in actual data.

It’s like comments below arguing about the reasoning levels not normalized to some metric (like cost, output token amount or duration) but just the labels or high, max, medium, etc. Those mean almost nothing even when comparing models based on the same pretrain (just compare GPT-5.4 to GPT-5.2), they mean less than nothing comparing different labs releases.

WarmWash an hour ago | parent [-]

It's not totally a mystery

https://arxiv.org/html/2604.24827v1

The short of it is by using hard facts knowledge that is difficult to compress, and then quizzing models on these facts and calibrating against a bunch of open models, you can kind of feel out the size of closed models.

duplessitous 6 hours ago | parent | prev [-]

> [...] shows an intelligence score of 59, the same as Opus 5 medium!

Nothing here is false, you are simply confused. You either didn't read what they wrote in its entirety or decided to reinterpret what they did write.

knollimar 2 hours ago | parent [-]

"Beating opus" is the false part, no?