| ▲ | Atlas-Finance: Evaluating AI Agents Inside a Bank(joinhandshake.com) | |
| 2 points by cjbarber 10 hours ago | 1 comments | ||
| ▲ | cjbarber 10 hours ago | parent [-] | |
I found this interesting. > We evaluate 11 frontier models (see Figure 1), all run within the OpenCode agentic harness. Claude Opus 5 performs the best, yet still only manages to pass 12.3% of tasks. Claude Fable 5.1 and GPT-6 Astra are close behind, but the other eight models have significantly lower pass rates. | ||