| ▲ | 2001zhaozhao 3 hours ago | |
It is probably from randomness. The benchmark tasks nowadays are so long that you can't really afford to run a large number of samples of them per model & effort combination | ||
| ▲ | artemisart 2 hours ago | parent [-] | |
No it's a mean of 5 runs. > We report FrontierCode’s overall score, a composite measure that grades each patch on blocking functional criteria (held-out unit tests) together with weighted code-quality rubric criteria, as mean@5. They don't explain more in the system card, I guess higher effort levels could loose points on the code quality / scope / style / maintainability stuff? | ||