Remix.run Logo
"Opus 5 is a really bad model"(twitter.com)
54 points by behnamoh 16 hours ago | 18 comments
dahdum 12 hours ago | parent | next [-]

I’ve been using Opus 4.8 heavily with Claude Code heavily and I haven’t noticed any problems at all with 5. It seems a bit more proactive (not as much pre-ban Fable though). My claude.md and agent prompts are both 〜20kb, and I use memories/rules/skills as well.

Agent to agent communication does appear to have improved considerably. I run the loops to the full context window and with proper working docs I don’t even notice a difference after compaction.

cjk 8 hours ago | parent | prev | next [-]

Anecdotally, I have also felt many of these same complaints.

In my case, I have a semi-autonomous loop where Claude writes some code, and uses `codex exec` to do an adversarial review. I had what should have been a trivial feature go for 13 rounds of review/fix before I stopped it, where each round was just flip-flopping the same logic back and forth to try to make the tests pass. Codex kept (correctly) re-identifying the issues Claude was flip-flopping on. Codex even suggested fixes that would have worked; Claude ignored them repeatedly. I never saw anything remotely this bad on Opus 4.8.

Additionally, I have a CLAUDE.md instruction to not silently defer anything, and ask me anytime it wants to do so. Opus 4.8 paid attention to this rule the overwhelming majority of the time. Opus 5 seemingly cannot be bothered.

I have tried updating my CLAUDE.md according to Anthropic’s recommendations for Opus 5, but it doesn’t seem to have made any difference.

K0balt 14 hours ago | parent | prev | next [-]

My experience with opus 5 has surfaced many of the complaints listed it the tweet. It seems to be slightly brighter than 4.8, but the laziness more than makes up for it.

docjay 2 hours ago | parent | prev | next [-]

Honest question: Does a shocking percentage of the population not use capitalization and punctuation anymore? Do people type like that in every situation now?

It seems like every time I see a screenshot of someone’s session it’s full of “u fix yet” or “why didnt you fix it” prompts. I thought I was only seeing posts from those people because those people are the ones that get their repo wiped by AI, but at this point it seems I’m the last of a dying breed or something. I even saw an OpenAI model launch that used ‘drunk text message’ formatting in the examples.

Edit: someone downvoted this comment, which I can only guess is because they’re used to people that use “honest question” as a prefix for a rhetorical or insulting statement dressed up like a question. That’s not what I’m doing. I don’t have the capacity to care how other people interact with their tools, it doesn’t affect me at all. I’m truly asking because I see it a lot lately, but not in my circle, so perhaps I’m a generation removed from it. My theory is that I learned to type on a keyboard, where autocorrect and capitalization was performed manually, but later generations learned primarily on a phone that handled all that for them. Perhaps switching to a keyboard to use Claude didn’t naturally engage the need to capitalize letters and such.

Grimblewald 10 hours ago | parent | prev | next [-]

mirrors my own findings as well. It royally bungles projects in a way that make git a bigger godsend than it should be (evem though it is).

I really miss the 4.5 era, that was a magic time.

steveharman 2 hours ago | parent [-]

Can't you just specify claude-opus-4-5 with the /model command and go back to that era?

Bawoosette 14 hours ago | parent | prev | next [-]

Most of this seems to be related to harness issues, specifically prompts provided by Claude Code. The upside of this is that it is likely cheap and easy for Anthropic to fix the issues by revising their approach to prompting their new models.

xyzsparetimexyz 10 hours ago | parent | prev | next [-]

Is it a really bad model in the same way that idk, Gemma 1.1 is a bad model? Stupid hyperbole

Frannky 9 hours ago | parent | prev | next [-]

I think they're doing something wrong. I have access to better models, but I just use 4.6 and 4.8. Fable and Opus 5 burn tokens too fast for me. I also switched to a cheaper plan to work less and to experiment with OpenCode and open models when I hit the 5h limits. I'm doing this partly because I feel they've already started the enshittification of the old models, and I want to be able to just stop paying for the subscription. I only use Claude via subscription plan—I've never needed to pay the insane API pricing, and I'm able to do everything with OpenRouter and cheap Chinese models (DeepSeek Pro, Flash, MiMo 2.5 Pro for reasoning, and MiniMax M3 for agents). I'm just still using them for coding, but I'm not sure for how long—I haven't nailed down an open harness with the new models that's cheap, good, and fast. Open to suggestions and ideas.

Schiendelman 6 hours ago | parent [-]

What's your project structure like? How many MD files do you have, how big is your backlog, your memory, how much is it loading at the beginning?

Cider9986 10 hours ago | parent | prev | next [-]

Was stunningly smart for the one question I asked it and it was on arena.ai so I didn't know which it was until voting for it.

13 hours ago | parent | prev | next [-]
[deleted]
dude250711 11 hours ago | parent | prev | next [-]

There had been a big pressure to increment the number.

It's not like users can prove anything about a remote black box anyway.

dgellow 12 hours ago | parent | prev [-]

Could we replace the link with https://xcancel.com/i/article/2081697911847481502 ? Faster to load, doesn’t nag you to sign up or install an app

orphea 11 hours ago | parent | next [-]

Usually the link remains original but someone helpful posts a comment with a link to xcancel or archive.org.

eddyg 11 hours ago | parent | prev | next [-]

The guidelines⁽¹⁾ make it clear the OP did the correct thing: ”Please submit the original source.”

⁽¹⁾ https://news.ycombinator.com/newsguidelines.html

Cider9986 10 hours ago | parent | prev | next [-]

Recently the xcancel captcha changed and now I prefer nitter.net. seems the same but no captcha.

fragmede 12 hours ago | parent | prev [-]

email hn@ycombinator.com