| ▲ | seizethecheese 2 hours ago | |||||||
> With Fable 5, Strands harness cost 77% less than Claude Code and scored higher on Terminal Bench 2.1. Terminal Bench 2.1 is saturated. Many token saving techniques would save money and score basically the same running Fable 5 against Terminal Bench 2.1. (They claim a better score but don’t say how much better. I’d bet my favorite hat that it’s not statistically significant.) This is at least the fourth time I’ve seen a project hit front page with a “save money with same score on saturated benchmark” claim. | ||||||||
| ▲ | strandstan an hour ago | parent [-] | |||||||
The scores are in blog post's bar chart. For Terminal Bench 2.1, Strands harness (Fable 5) scored 69.7 while Claude Code (Fable 5) scored 61.8. This is on high effort. I hear you tho about saturation. We're working on a follow-up deep dive post with more harnesses, so could look into Terminal Bench 4.0? | ||||||||
| ||||||||