| ▲ | nojs 5 hours ago | |||||||
We really need better harness benchmarks. It seems there's no reliable source that benchmarks the main harnesses against all open source models. I also wish the discussion around Pi did not always use cost/token count as the metric. It's amazingly token efficient, but how does it stack up again opencode and others if you don't care about token count? My experience is that the harness is mainly polish preventing failed tool calls, bad edits, stuff like that, but doesn't make much difference to the overall "intelligence". But that opencode seems slightly more robust against stupid errors than out of the box Pi due to the additional context it forces through every thread. | ||||||||
| ▲ | calgoo 3 hours ago | parent | next [-] | |||||||
I really disliked how opencode works IMO; the harness tries to do to much in my mind. Switching to Pi was a breath of fresh air for me, and I even use hax for some of my local needs where i dont want to have the giant pile of fertilizer that is NPM or PIP installed. The harness becomes more and more important, the smaller the model is as you need to offload context management as well as memory to the harness. The big models basically just need a bash prompt tooling and you let the model manage everything inside its own context. | ||||||||
| ||||||||
| ▲ | sn0n 4 hours ago | parent | prev | next [-] | |||||||
A good benchmark would require a decent number of smaller scoped one off tasks to larger multi step refactors, and also one shot full project of simple to complex varieties. In addition to a series of “conversational” ambiguity filled one-liners. | ||||||||
| ▲ | sandeepkd 2 hours ago | parent | prev | next [-] | |||||||
Harness and benchmark for the harness feels like a chicken and egg problem. The harness is to optimize the interaction results with the models. Any benchmark for harness has to focus on the goals that the harness was trying to optimize for unless we are only focussing on generic harnesses. At this point when all the models have been trained on all available data with the similar algorithm, 1. either you get more data which is not feasible, 2. or get a better algorithm - a possibility , 3. or write a more targeted harness. Harnesses for legal, medicine and all are the ones which are getting focus for this reason. Writing benchmarks for these targeted harnesses would be a catching task | ||||||||
| ▲ | leemysw 4 hours ago | parent | prev | next [-] | |||||||
[flagged] | ||||||||
| ▲ | general_reveal 2 hours ago | parent | prev [-] | |||||||
Any harness will always be privately and secretly shaped and created by those selling models, especially coding models. It is literally a “selling point”, and unless the government steps in to oversee the tests like in the car industry, then there is absolutely no way lizard satanists like Altman and Musk are going to exercise their native ethical traits (“native” loosely assumes they procured it divinely and quite recently, because, we simply haven’t observed it prior). Short of that, these tests will be fabricated, a lot, for money. Tell a horny monkey not to jerk off. How the fuck … would that even be possible? God, only God can stop this godless train. There is no sincere discussion to be had here. HN has been a cesspool for marketing and it reeks in here lately. Edit: I am not punching down, it’s gotta be crooks from those companies down-voting. | ||||||||