Remix.run Logo
Aldipower 3 hours ago

I would say "Elected errors _in_ Claude Opus 5" wouldn't be incorrect either.. Opus 5 isn't very reliable for coding and introduces a lot of regressions every single time I use it. Do you have the same experiences?

flaburgan 2 hours ago | parent | next [-]

I also found it to forget some obvious cases in a quite simple flow (validate the email address of a user who register), that surprised me a lot. Maybe it's because I got used to Fable? But I am quite sure Opus 4.8 wouldn't have make this mistake. If I had time I would try the same prompt with it to see. Anyway, back on 100% Fable for me.

egeozcan 3 hours ago | parent | prev | next [-]

I don't know what I could be doing differently to you but I found Opus 5 to be more reliable than even myself at times. Maybe your stack is unusual or you have conflicting commands in your prompts vs CLAUDE.md (that really confuses it)? It could be anything but this huge error bar in delivered quality is one of the biggest issues with LLMs.

jarym 2 hours ago | parent | prev | next [-]

Getting to grips with each new model does require some tweaking and experimentation. So far I've found Opus 5 to repeatedly pause its work and give me some seemingly randomly invented decisions to make.

jdthedisciple 2 hours ago | parent [-]

> give me some seemingly randomly invented decisions to make

Any examples?

benjiro29 3 hours ago | parent | prev | next [-]

Opus 5 isn't very reliable for coding and introduces a lot of regressions every single time I use it.

And GPT 5.6 Sol over engineers just about everything. No LLM is perfect, its about learning the issues with each LLM and figuring out if you can live with it. Knowledge means that you can anticipate if it tries to pull something funny, and harness it against that behavior.

Aldipower 3 hours ago | parent | next [-]

Sure, but I am a long time Opus user 4.5,4.6,4.7,4.8 and I wonder what's wrong with 5?

pimeys 2 hours ago | parent [-]

I remember when 4.7 and 4.8 were released and people were asking what's wrong with them and 4.6 is the best.

But yes, I also think it's not the greatest model for programming. On the other hand, for agentic tasks that are not programming related it's hard to beat Opus 4.8. It can try different things and pivot even when the user is not great with prompting. 5.0 seems to not be worse, but definitely wastes more tokens and costs more.

someguyiguess 2 hours ago | parent [-]

4.6 was better in some way that I can’t put my finger on. None of the models since have been able to reproduce its quality of output for me.

cyanydeez 3 hours ago | parent | prev [-]

this would be _Great_ advice if you owned your own LLM and your knowledge was trapped in Amber because you were satisfied.

It's horrible advice given what we've seen consistent: changing alignments, changing guardrails, changing system prompts, changing inference priorities, etc.

Anyone who relies on these for their work product is chaining themselves to a matrix multiple of indetermintism.

Saline9515 2 hours ago | parent | prev | next [-]

Same here, also it lies often to me or implements something else that what was planned. It feels quite strange to see it say casually "I didn't tell you the full truth on X" when I notice the issues. At the same time, maybe it is more honest?

bearjaws 3 hours ago | parent | prev | next [-]

I do not experience any regressions, I don't really notice much difference either.

trentor 3 hours ago | parent | prev | next [-]

Maybe it's my harness but I haven't seen it introducing regressions.

cbg0 3 hours ago | parent | prev [-]

Do you not have unit tests, or how does it introduce regressions? You can tell it how to run the test suite in CLAUDE.md

quaheezle 3 hours ago | parent | next [-]

Opus 5 tries to modify the unit tests as a cover to its own regressions - thinking its own logic is correct and the test must be wrongly specified

jvuygbbkuurx 2 hours ago | parent [-]

I found it eagerly reversing existing product decisions like changing a user given date into created date, since it thought creating something that is in the past is incorrect. This was not even related to the task at hand at all. I noticed this kind of stuff happens more with ultracode for some reason.

grim_io 2 hours ago | parent [-]

Any subagents-based workflow is prone to this, because of the fragmented context(by design).

Aldipower 3 hours ago | parent | prev | next [-]

I detect the regression already in planning with Opus 5, so I do not let Opus 5 implement anything. But it is a waste of time and tokens! Does planning with Opus 5 works out for you?

cbg0 2 hours ago | parent [-]

It sounds like you just need to correct the plan it lays out to avoid the regression? I'm just looking to debug with you, not defending the model. I've mostly used Opus 5 for code reviews & bugfixes.

Aldipower 2 hours ago | parent [-]

Yeah fine. I mean my plan wasn't to difficult. For example this morning I started with Opus 5 to tackle a problem. During planning at some point Opus 5 detected _8_ regressions in it's own planning, after I directed it towards those potential regressions. So, in this very moment now, Fable 5 implements code already and the planning before with Fable, done with the same instructions, was flawless and quick. And I am sure, Opus 4.8 would be flawless either. Same harness, same claude.md, etc..

cbg0 2 hours ago | parent [-]

Do you think it might be overthinking? Try it on medium effort.

Aldipower 2 hours ago | parent [-]

Probably, I ran it on xhigh. Worth a try.

simiones 3 hours ago | parent | prev [-]

It is a well known fact that projects with unit tests never have regressions.

cbg0 2 hours ago | parent | next [-]

I don't understand the need for the snarky comment, LLMs can run the test suite and avoid regressions.

simiones 2 hours ago | parent [-]

A change can introduce regressions in any large project even if 100% of unit tests pass. Unit tests test individual units, regressions can happen at many levels. Especially if we treat performance degradations as regressions.

cyphar 2 hours ago | parent | prev [-]

[dead]