| ▲ | strandstan an hour ago | |
The scores are in blog post's bar chart. For Terminal Bench 2.1, Strands harness (Fable 5) scored 69.7 while Claude Code (Fable 5) scored 61.8. This is on high effort. I hear you tho about saturation. We're working on a follow-up deep dive post with more harnesses, so could look into Terminal Bench 4.0? | ||
| ▲ | seizethecheese an hour ago | parent [-] | |
Then I'm really confused. Terminal Bench 2.1 scores on Artificial Analysis are like 90%. https://artificialanalysis.ai/evaluations/terminalbench-2-1 | ||