Remix.run Logo
traceroute66 2 hours ago

So TL;DR benchmarking in a completely non-reproducible manner ?

"Model X performed great, but we can't possibly tell you anything about the code it was looking at apart from it was a large code base from an unknown company".

So basically pinky-promise benchmarking ?

I'm not sure I follow the value here ?

kadoban 2 hours ago | parent | next [-]

If it builds up history and perceived reliability, this type of thing can be valuable. You're giving up transparency for it being harder to game.

traceroute66 2 hours ago | parent [-]

> You're giving up transparency for it being harder to game

But then if we take that argument to its natural extreme, surely it means people should take the marketing bullshit published in the 100-page system cards published by Anthropic & co as "valuable" too ?

kadoban 2 hours ago | parent [-]

I think you know that's basically nothing like this? The model cards have every incentive to be biased, this doesn't necessarily.

But even so, pretty much yes: companies that actually have reliable and accurate info in their releases get trusted more. It takes time because the default is to disbelieve info from biased sources, but it is possible to trust some of them more than others.

deepwoods 2 hours ago | parent | prev | next [-]

In theory, as long as all the models are doing the same thing with the same tools, it's at least useful to see how they stack up against each other right now. It might not be great to track progress over time, as it can get benchmaxxed or the underlying resources may become obsolete.

sigmar an hour ago | parent | prev | next [-]

Lots of private benchmarks already exist, where you have to trust the tester (ex Artificial Analysis, Arc-agi).

demibabs 2 hours ago | parent | prev [-]

Doesn’t it ultimately have to be this way, to prevent saturation?