| ▲ | goldenarm 6 hours ago | ||||||||||||||||||||||||||||||||||
Geometric mean of all these benchmarks : * GPT-5.6 Sol: 65.5 * Fable 5 (w/ fallback): 64.5 * Opus 5: 64.0 * DS-V4-Pro 0813: 62.5 * Kimi-K3: 62.3 * DS-V4-Flash 0731: 55.8 * GLM-5.2: 47.3 | |||||||||||||||||||||||||||||||||||
| ▲ | svachalek 5 hours ago | parent [-] | ||||||||||||||||||||||||||||||||||
Maybe it's me but I don't see how DS Flash is better than GLM at all, much less by a huge gap. I'd probably protest less against Fable and Opus being put at the same level than many would, but there's no denying the two models are a very different experience from each other. I guess where I'm going is no one should pick a model by the benchmarks. | |||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||