| ▲ | nightpool 4 hours ago | ||||||||||||||||||||||||||||
You're comparing the "score percentage" (e.g. out of the total number of partial points available, how many did the agent achieve) to the "completion percentage" (how many tasks does the model score 100% on). The paper says "Claude Opus 4.8 with maximum thinking and batched tool calls scores best but still completes only 20.6% of tasks at a 54.8% partial score", which is ~the same number that Anthropic reports here (55.7 vs 54.8). That is—the agent scored 100% on 20% of tasks, but on average it got 54% of the "score" awarded in the exam. One number reflects partial progress, the other one doesn't. The authors of the benchmark prefer you to look at the lower number (because they want to show their benchmark as capturing useful gaps in capabilities and with a lot of room for improvement), the authors of the models want you to look at the higher number (because they want you to think of their models as capable) | |||||||||||||||||||||||||||||
| ▲ | tadfisher 4 hours ago | parent [-] | ||||||||||||||||||||||||||||
In what world is 55.7 the same number as 54.8? What variance is acceptable to publish without a retraction? | |||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||