| ▲ | HarHarVeryFunny 10 hours ago | |
There's a very interesting benchmark comparison below of Opus 4.7 run under three different harnesses : OpenCode, Cursor and Claude Code where it's not very close at all and Opus's native harness, Claude Code, performs worst of all three. The pass@1 scores are 50/45/40 for OpenCode/Cursor/Claude Code respectively. https://artificialanalysis.ai/agents/coding-agents#harness-c... I've seen other benchmarks where Pi also outperforms Claude Code both in model performance and in much reduced token usage. | ||