Remix.run Logo
Roark66 an hour ago

It is useful to compare like for like.

Currently if my hypothesis about frontier labs doing creative tricks between the model and the client is true (and the results seem to favour it so far) the benchmarks are giving us an artificially lowered results for open weights models.

I have yet to test opus/sonet via my proxy. If Qwen gets 10% better and Opus stays the same that suggests one if two things: - either opus doesn't need it - or it's already done behind the scenes.

sanderjd an hour ago | parent [-]

In my view, the useful like for like comparison is to the entire system that people actually use. Nobody uses a "naked model", so what is the point of these benchmarks that use them in that way?

I think the benchmarks should be trying to use realistic harnesses for both proprietary and open weight models.