Remix.run Logo
jchw 3 hours ago

Almost 10 models are passing the benchmark >95% - isn't that... substantially overly saturated?

I'm also really skeptical of benchmarks that place any Haiku model very high. I've been thoroughly unimpressed with Haiku and my opinion has been that you basically shouldn't use it. Yet here, it ties Kimi K3. HMMMM.

"Rubric Quality" seems a bit more realistic than "Pass Rate", but it is apparently judged by Fable 5...

This entire write-up is also obviously clearly very heavily AI-assisted, which doesn't help matters any.

sambusa_123 2 hours ago | parent | next [-]

Just read through some of the code "benchmarks", and I see why: https://reinvently.co.uk/tools/ed-o-meter/tests/

Most are extremely trivial tasks. I would be surprised if a model from 2 years ago failed these...

nylonstrung an hour ago | parent [-]

I think most of them would have actually failed, it's only recently that models were any good at using tool calls and harnesses after they started post-training for that

rfgplk 2 hours ago | parent | prev | next [-]

> Almost 10 models are passing the benchmark >95% - isn't that... substantially overly saturated?

The only true benchmark for any of these models that I've discovered isn't if they can pass precanned SWE tests, but rather can they create something novel? This isn't even too difficult to test, just give it a seemingly impossible task let it spin and see where it ends up.

finaard 2 hours ago | parent | prev | next [-]

Interesting - when Kimi K2.6 came out I switched over from Anthropic models, with at that time comparable to better results for me. I was using Anthropic via API, heavier months were roughly $400 worth of Anthropic tokens - I can get the same thing done via a $100 ollama subscription.

zuzululu 21 minutes ago | parent | prev [-]

yeah that haiku really diminishes the claims behind the benchmarks. luna-max is significantly cheap and it is a strong performer but i dont see it on the benchmarks.

I just get a feeling this site started with an intent to elevate Chinese models above the rest so wouldn't be surprised if the whole prompt sail was set to that tune