Remix.run Logo
kennywinker 5 hours ago

Does it help with benchmarks? Are you saying there are examples of benchmarks where the models have solved the problem by cheating?

well_ackshually 5 hours ago | parent [-]

Hundreds, at this point? Every benchmark is flawed as shit, written by clowns. DeepSWE, They're given the full git history (the solution is in it), others don't even bother to verify if the code is the right one and just the output, they've modified the test harnesses, injected code to make all tests pass, etc. The entire benchmark galaxy is just clowns propping eachother up and are regularly talking with the big AI labs.