Remix.run Logo
egeres 6 hours ago

It feels suspicious that MiMo-V2.6 Pro gets 46 in de index while DeepSeek-V4.1 (https://artificialanalysis.ai/models/deepseek-v4-1-flash) gets 39. According to the appendix at the bottom of https://mimo.xiaomi.com/mimo-v2-6 the deepseek model sometimes surpasses mimo and it's not so far behind in capabilities. A week ago opus 5 appeared 1 points ahead of fable 5 despite fable being a much smarter model (this has been corrected already)

SyneRyder 4 hours ago | parent | next [-]

The main AA benchmark keeps changing, and had to be radically changed when Astra came out and showed zero improvement over GPT 5.6 Sol in their benchmark. Opus 5 is still 1 point ahead of Fable 5.0 on the index, if you manually add Fable 5.0 back into the list, so it hasn't actually been "corrected". It's only Fable 5.1 that is shown as ahead of Opus 5.

The AA benchmark is a weighted average of other benchmarks and some internal ones. I think the difficult part is finding benchmarks that reflect your own use of the models.

seahorseemoji 2 hours ago | parent | next [-]

The way Artificial Analysis keeps changing their weights feels kind of like deciding who the winner should be and making the weights reflect that. They’ve been changing their weights to add more weight to improved long-running agentic capabilities, but doing so means they’re reducing the relative importance of world knowledge and of writing ability.

I’ll grant that maybe world knowledge isn’t that important for these models. But writing ability is important for human understanding, and I think the weird turns of phrase and word choices reflect the labs’ underweighting of the importance of human understanding.

sipjca 2 minutes ago | parent [-]

I mean artificial is in their name....

yt1998 36 minutes ago | parent | prev [-]

[dead]

GodelNumbering 4 hours ago | parent | prev [-]

> It feels suspicious that MiMo-V2.6 Pro gets 46 in de index while DeepSeek-V4.1 gets 39.

Why?

big-chungus4 an hour ago | parent [-]

> W

W what?