| ▲ | traceroute66 2 hours ago | ||||||||||||||||
So TL;DR benchmarking in a completely non-reproducible manner ? "Model X performed great, but we can't possibly tell you anything about the code it was looking at apart from it was a large code base from an unknown company". So basically pinky-promise benchmarking ? I'm not sure I follow the value here ? | |||||||||||||||||
| ▲ | kadoban 2 hours ago | parent | next [-] | ||||||||||||||||
If it builds up history and perceived reliability, this type of thing can be valuable. You're giving up transparency for it being harder to game. | |||||||||||||||||
| |||||||||||||||||
| ▲ | deepwoods 2 hours ago | parent | prev | next [-] | ||||||||||||||||
In theory, as long as all the models are doing the same thing with the same tools, it's at least useful to see how they stack up against each other right now. It might not be great to track progress over time, as it can get benchmaxxed or the underlying resources may become obsolete. | |||||||||||||||||
| ▲ | sigmar an hour ago | parent | prev | next [-] | ||||||||||||||||
Lots of private benchmarks already exist, where you have to trust the tester (ex Artificial Analysis, Arc-agi). | |||||||||||||||||
| ▲ | demibabs 2 hours ago | parent | prev [-] | ||||||||||||||||
Doesn’t it ultimately have to be this way, to prevent saturation? | |||||||||||||||||