Remix.run Logo
Fordec 2 hours ago

This all strikes me as an effort to move tailoring the harness out of the easily transferable .md file into specific Anthropic tooling to increase lock in.

I've been running Opus 5 today and it's already done accidental deletions, made far more mistakes and worked around deliberate hook controls than previous Opus versions combined. Also it looks like token usage is up as it fails at the task the first time around much more frequently than 4.8.

frio an hour ago | parent | next [-]

I’m not excited about using Opus 5, mainly because the way that I work atm — essentially peer programming — means I sandbox the agents and work with them closely. Opus 4.x encounters the sandbox and moves on with its day; Fable becomes increasingly fixated on it and does less and less of the actual task, focussing more and more on the limit it reached. I worry that, from your description, Opus 5 will do the same.

Fordec an hour ago | parent [-]

It does feel a bit smarter, but it seems to be also "too clever by half" and its not ignoring the rules, it's rejecting them and finding work arounds. It sticks to the word of the law, while rebelling against the spirit of the law.

One example is to get around a git --checkout usage ban, it CD'd to another folder first and back to bypass the regex in the hook.

nextaccountic 20 minutes ago | parent [-]

This is called misalignment

We may some day find out that the smarter the model is, the hard is to align it properly

vidarh an hour ago | parent | prev | next [-]

I have a document generation task that I used to run with 4.8. This morning after it switched to 5, the documents were consistently 30%-40% longer for the same prompt... Not evaluated whether they are actually better or worse yet, but what was interesting was how consistently more verbose it was.

wren6991 2 hours ago | parent | prev [-]

I think the type of persistence rewarded by benchmarks may be misaligned with instruction following