Remix.run Logo
dsiegel2275 3 hours ago

The focus of the study was the different harness approaches and how they scale across model sizes. The fact that they used any particular set of models is irrelevant.

Tycho 5 minutes ago | parent | next [-]

But the behaviour of the system can totally change under different scales.

comparing x10 to x100 doesn’t necessary inform you about x100_000 to x1_000_000

klooney 2 hours ago | parent | prev | next [-]

I'm not totally convinced that models are fungible, the claudes/gpts/Gemini all have pretty individual feels when you're working with them. I wouldn't be surprised if the approaches don't scale or even work the same in a poly model setup

paimapi an hour ago | parent [-]

that's anecdotal though, right? your subjective feeling of how a model responds to you will greatly influence how you 'feel' about a model and its performance in the same way that a co-worker who you get along with will fuck something up and you'll be more forgiving than when you work with a too-verbose, mansplainer of a co-worker who fucks up

I think until we have actual repeated-use measurements tracked over time (eg consistent prompts used to do the same tasks, count number of hallucinations and errors and bugs over a long period of time) you won't really have any idea of which model is better

I also think of it like a car - some just feel better to drive even if they are materially worse in other measures. until you start measuring the metrics important to you (eg MPG and cost of maintenance over a long period), you have no idea which car is actually better suited for you. and the fact that you can only do so with a limited number of cars (or hours available to work, or money to burn on tokens) means there's no true measure approaching objectivity

virgil_disgr4ce an hour ago | parent [-]

Anecdotal or subjective don't mean "wrong." I would 100% agree that claude and chatgpt have different 'styles.' They do have their own patterns, and those patterns are distinguishable.

hiddencost 3 hours ago | parent | prev | next [-]

Nope. Sorry. Not how this works.

bjelkeman-again 3 hours ago | parent [-]

How does it work then?

belowavgiq 2 hours ago | parent [-]

I'm going to cherry pick one example where newer models are noticeably improving at least in my experience.

What is a noticeable improvement with something that struggles to read a message longer than 200 characters without missing information in the middle, may be a 0.000000001% improvement with a model that... almost never misses info in the first place.

shermantanktop 3 hours ago | parent | prev [-]

Agree. Harnesses are effective because they interact with the underlying model effectively. If the latest models were fundamentally different, excluding them would be a miss. But I don’t think they are, at least not in ways that would affect these observations.

svachalek 3 hours ago | parent [-]

I think the confounding issue is that by now, millions of sessions of Claude Code and Codex are now in the training set for these models. So they have been trained to work the way these harnesses are configured, and at least in the case of Claude Code the harness itself is greatly stripped down because the model has absorbed it.