| ▲ | sosodev 5 hours ago |
| Because AA Coding "Index" consists only of two benchmarks (Terminal-Bench v2.1, SciCode) and generally fails to be meaningfully representative of agentic coding capabilities. |
|
| ▲ | Alifatisk 5 hours ago | parent [-] |
| Whats a better option for AA Coding Index? |
| |
| ▲ | WASDx 4 hours ago | parent [-] | | DeepSWE and FrontierCode are more realistic if you read up on what they actually measure. But the most realistic is to try it yourself. Benchmarks can only vaguely represent typical usage, and how you judge the result. Giving the same real task you have to a few models will make you understand them better than chasing benchmarks. |
|