Remix.run Logo
onomojo an hour ago

Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.

Fordec 14 minutes ago | parent | next [-]

Yeah, I've dropped back to 4.8 entirely for the remainder of this billing cycle. I'm going to be seriously looking into Qwen adoption and harness migration options over the course of August.

cromka an hour ago | parent | prev | next [-]

Agreed, it's extremely frustrating. It's the only model that actually makes me curse when talking to it, even knowing how counterproductive it is.

copperx an hour ago | parent | prev | next [-]

I'm dumbfounded to see Opus 5 making SO MANY mistakes in coding simple stuff. Most times, Fable 5 comes out to be cheaper because it nails so many things much quicker than Opus 5.

garciasn an hour ago | parent | next [-]

I have Fable plan and Opus implement. I haven't had any major issues working this way; however, Opus does seem plain fucking stupid compared to what I experienced with Sonnet previously.

aenis 40 minutes ago | parent | next [-]

I do the same, and generally have good results, but it does stupid things with gusto.

I'd open a blog with "weird things Opus did". Today it launched a swarm of cpu-hogging processes to test if the widget showing machine and I/O load is rendering nicely and correctly. The test went fine, but it was no longer able to kill those processes since they were really effectively hogging the CPU in various ways - being diligent, some of them were hogging CPU, some were murdering the SSD, some were pounding on the network adapters. Took me 30 mins to recover the machine to a working state without killing the meaningful, messy, in-flight sessions i had going on on other projects.

petesergeant 32 minutes ago | parent | prev [-]

> however, Opus does seem plain fucking stupid

Infuriatingly so, in a way I don't remember Opus 4.8 being, but maybe I've just been ruined by Fable 5.

hbn 14 minutes ago | parent [-]

I bought my first LLM subscription with Claude right before they gave access to Fable 5.

I got so used to it, when they finally pulled access for me and I had to go back to Opus I felt like I was working with my hands tied.

I finally know what those women with AI boyfriends felt like when their app updated and it won't dirty talk with them anymore.

usef- 37 minutes ago | parent | prev [-]

Weird how different people's experiences are. If it's making simple mistakes something must be wrong in your setup/context I assume? It's been solid for me, beyond the usual LLMisms that all models have. But I keep context pretty minimal.

cromka 20 minutes ago | parent | next [-]

Statements like this typically come from working on the same setup and context using different models. I actually have that very experience now; I work on something security-adjacent so Fable often drops out, at which point Opus behaves like its lobotomized half-sibling. Pardon me the language, but I can't find a better example to be honest.

nimonian 11 minutes ago | parent | prev [-]

Agreed. Opus 5 is doing just fine, slightly better than 4.8. It's personality is insufferable, but I find myself catching fewer problems at code review. It generally understands my conventions and isn't so eager to accrue tech debt.

visarga an hour ago | parent | prev | next [-]

Sent to solve one task, came back with half of it solved and 2 more problems.

capnjazz 38 minutes ago | parent [-]

"One thing worth your attention", "Two things worth knowing", "One thing to eyeball"

greenchair a minute ago | parent [-]

This is driving me crazy. opus 4.8 did not do this to me not (at least during pre-5.0 timeframe). Feels like the new cycle is one step forward, two steps back.

enraged_camel 40 minutes ago | parent | prev | next [-]

It's my daily driver. I like it and find it noticeably better than Opus 4.8.

After I started reading complaints about Opus 5, I gave Fable the task of evaluating a bunch of code Opus 4.8 had written and compare it to Opus 5's code. Fable ran a dynamic workflow and the scores came back 15-20% higher for Opus 5's code in terms of quality, correctness and readability/conciseness. I did not tell Fable which Opus wrote which code, and I turned off memory as well to ensure there was no pollution from that angle.

My only complaint is that Opus 5's prose is annoying as hell. I wrote a custom skill for it for concise debriefs and it has been working pretty well for me.

nomel 41 minutes ago | parent | prev | next [-]

What's the clear best, that you see?

drschwabe 28 minutes ago | parent | next [-]

GPT 5.6 Sol

petesergeant 38 minutes ago | parent | prev [-]

Fable 5

fellowniusmonk 25 minutes ago | parent | prev | next [-]

I have some internal tests I use for areas where one particular solution/paradigm is dominant but worse.

Opus 4.6 is the last model that's actually useful and can "adjust" its perspective to use the newer & better solution.

Where Opus 4.8-5 has over fit training on worse/older but "dominant" solutions it refuses to adjust.

Not only does this create an existential threat to adopting progress but it also means that if you have a code base that has rare but real world tradeoff the newest versions of Opus 4.7, 4.8 and 5 are worse than useless and become a major dev timesink.

logicchains an hour ago | parent | prev | next [-]

"As you requested, I've finished task X. Honestly, task X turned out to require task Y, which I haven't actually done. Task Y is the next step if you'd like to continue along this route."

pornel 29 minutes ago | parent [-]

This is the hard-won load-bearing quote.

bontaq an hour ago | parent | prev [-]

It's an infuriating model