| ▲ | hadlock 4 hours ago | |||||||
When it comes to quality of outcome, since at least Feburary, the harness has almost equal, if not more weight than the model itself. It's no longer "which model is the best?" it's "which model + harness is the best?" I get drastically different tool call failure rates using Claude SDK vs OpenCode using Qwen 3.6 models | ||||||||
| ▲ | RideOnTime22 11 minutes ago | parent | next [-] | |||||||
Every other week it's a new "X didn't matter, until Y date" without any hard quantitative claims. It's crazy how over the past years a field originating from math ends up succumbing to subjective feels. | ||||||||
| ▲ | JLO64 4 hours ago | parent | prev | next [-] | |||||||
It's worth nothing that recent Claude models seem to have gotten worse at tool calling outside of Claude Code and the SDK: https://lucumr.pocoo.org/2026/7/4/better-models-worse-tools/ | ||||||||
| ▲ | KronisLV 4 hours ago | parent | prev | next [-] | |||||||
> the harness has almost equal, if not more weight than the model itself This feels like a horrible failing of the models to generalize, then - both basic and intermediate tasks should be possible to do with Claude Code, OpenCode, Pi, ZCode, Kimi Code, Dirac and tbh any other mainstream or even slightly niche harness. Not doubting the claim itself, there's a reason why good benchmarks include the harness. | ||||||||
| ||||||||
| ▲ | HDBaseT 29 minutes ago | parent | prev | next [-] | |||||||
Yeah this is a complete lie. You can use effectively any harness and get good results. Harnesses are mostly placebo. | ||||||||
| ▲ | davidlt 4 hours ago | parent | prev | next [-] | |||||||
I just wanted to emphasize this. Harness is a big part of how things perform thus usually it's harness + model co-design that's important. | ||||||||
| ▲ | azinman2 4 hours ago | parent | prev [-] | |||||||
Which works better for you? | ||||||||