Remix.run Logo
chewz a day ago

> If by Opus you mean Opus 4 and not Opus 4.8, then sure

I meant Opus 4.8 which is rather dumb and ineffective in coding harness, especially with higher thinking levels.

porksoda a day ago | parent | next [-]

My experience was so much different to this, that I have the unfortunate impression that you're shilling. It really was not a capable model, it felt like the old oai models back when we were all excited but couldn't actually trust them even in the littlest ways. What harness were you using, did you do any work to make it better? What was I doing wrong? I just pointed opencode at it, with a pretty simple (large-ish) data cleaning project.

VladVladikoff a day ago | parent | next [-]

It’s so fascinating watching people on here bicker about models like fine wines. Wild times we live in. Wild times.

nozzlegear a day ago | parent [-]

It's not a frontier LLM if it's not made in the Silicon Valley region of California, otherwise it's just a sparkling LLM.

Planktonne a day ago | parent | prev | next [-]

> My experience was so much different to this, that I have the unfortunate impression that you're shilling

I think you might both just be reading way too much into one-off random experiences that you've decided are evidence of significant and stable capability.

chewz a day ago | parent | prev [-]

Have you actually used Opus 4.8 in Claude Code? It takes way too long to do any practical task on higher thinking levels due to over-engineering. And I am not the only one complaining. Lots of people downgrade to Opus 4.6 exactly for this reason.

Opus 4.8 training works well for agentic work. Not for code harness.

EDIT:

```

stronger on coding and raw capability but can be more argumentative, verbose, and costly.

Reliability and instruction-following Many users say 4.6 felt more reliable and followed instructions better. "With 4.6, when I tell it something, it actually remembers the spirit of what I asked for and keeps applying it."

Others report 4.8 drifts from preferences and can be frustrating to control. "I still find myself getting frustrated when it ignores preferences and drifts from instructions"

Some people find 4.7/4.8 push back more and act more adversarial than 4.6. "The biggest complaint against 4.8 is that it is argumentative and "pushes back" constantly"

Coding quality and capability Several users praise 4.8’s coding strength and thoroughness. "4.8 is technically impressive, especially for coding"

Other reports say 4.6 could be better for certain coding workflows and breaks less. "4.6 still >> 4.8 for anyone else as well? Maybe I'm in the minority, but for my use cases Opus 4.6 is still better than"

Some recommend mixing models: use 4.8 for key tasks and 4.6 for general work to save tokens. "What I do is... use 4.8 for key moments, and for everything else 4.6"

Cost, speed and token behavior Users note 4.8 often uses more tokens and can feel slower because it “thinks” more. "4.8 is much more cautious, and as a result - slower. It checks everything, thinks for a long time etc."

```

[https://www.reddit.com/answers/601770d4-4059-478d-aa52-b445c...]

anonzzzies a day ago | parent | next [-]

Works extremely well for us. I never know what other people are doing when we read these stories.

rapind a day ago | parent [-]

Yeah I found 4.7 and 4.8 to be downgrades from 4.6. I don't know if it was the model or just Anthropic's scaling issues though TBH. I found working with the Claude Code max (20x) sub would work awesome in between new models. A week before up to 2 weeks after, it would go to crap, dumber, slower, outages, harness churn, etc.

I'm finding the same with ChatGPT recently since the 5.6 release. Not as bad though, but sluggishness at times, harness churn (creating bugs and crashed), and occasional availability issues that cause me to downgrade to 5.5.

It's gotten to the point where I dread a new model release from these companies because it's guaranteed to be disruptive! I assume the pay per use API is less impacted.

sunaookami a day ago | parent | prev [-]

Reddit is the absolute worst way to find real experience. Opus 4.8 is a very capable model and no chinese model can outdo it.

Narciss a day ago | parent | prev | next [-]

I can’t believe that anyone would actually think this. This

a day ago | parent | prev | next [-]
[deleted]
mattmanser a day ago | parent | prev [-]

Comments like this boggle my mind.

The model which everyone else raves about and is wildly successful with legions of programmers virtually demanding access while abandoning ChatGPT and Copilot in droves, is rather dumb?

Have you considered that it's more likely that you're doing something wrong?

piguin a day ago | parent | next [-]

My own experience is that the vast majority of programmers have experience with one model and maybe some short usage of earlier models from a competing choice but want to be using the model with the highest popularity and reputation. I've worked with people who actually had to test multiple choices for their team who didn't understand why they were pressured to select Claude for programmer morale.

AussieWog93 a day ago | parent | next [-]

I went from cycling between models all the time in Cursor (some would randomly be better at certain tasks than others) to just going pure Opus 4.5 when that came out - it was so far ahead of anything else at the time.

Interestingly with Fable vs GPT-5.6 I think they've lost their lead a bit. I'm finding Fable can't do certain work that 5.6 Sol Ultra can - especially when it comes to webpage design.

Grok 4.5 was fast but made mistakes that GPT/Fable just don't.

I'm curious to try Kimi.

WarmWash a day ago | parent | prev [-]

When iOS users are given an Android phone, they complain about how awful of an OS it is.

In reality, they just aren't used to it.

Levitz a day ago | parent [-]

What exactly is there to get used to? Do people really use each model very differently? There's a learning curve as in everything, but there is nothing that comes to mind for me when I use codex as opposed to claude

ivewonyoung a day ago | parent | prev | next [-]

> is wildly successful with legions of programmers virtually demanding access while abandoning ChatGPT and Copilot in droves

Do you have a source for that? Codex went from 5 million users to 9 million users in the past few weeks since GPT 5.6 released. It was so popular that Claude was forced to extend Fable access by a week and then permanently for some plans.

gigatexal a day ago | parent | prev [-]

Exactly what I’m thinking, too.