| ▲ | spongebobstoes an hour ago | |
this is not a good measure of current model capability. we need to test agents in harnesses, not models with a single prompt test Codex, not Sol. test Claude code, not Opus | ||
| ▲ | ChrisLTD 22 minutes ago | parent [-] | |
There are other benchmarks for that | ||