| ▲ | porksoda a day ago | ||||||||||||||||||||||
My experience was so much different to this, that I have the unfortunate impression that you're shilling. It really was not a capable model, it felt like the old oai models back when we were all excited but couldn't actually trust them even in the littlest ways. What harness were you using, did you do any work to make it better? What was I doing wrong? I just pointed opencode at it, with a pretty simple (large-ish) data cleaning project. | |||||||||||||||||||||||
| ▲ | VladVladikoff a day ago | parent | next [-] | ||||||||||||||||||||||
It’s so fascinating watching people on here bicker about models like fine wines. Wild times we live in. Wild times. | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | Planktonne a day ago | parent | prev | next [-] | ||||||||||||||||||||||
> My experience was so much different to this, that I have the unfortunate impression that you're shilling I think you might both just be reading way too much into one-off random experiences that you've decided are evidence of significant and stable capability. | |||||||||||||||||||||||
| ▲ | chewz a day ago | parent | prev [-] | ||||||||||||||||||||||
Have you actually used Opus 4.8 in Claude Code? It takes way too long to do any practical task on higher thinking levels due to over-engineering. And I am not the only one complaining. Lots of people downgrade to Opus 4.6 exactly for this reason. Opus 4.8 training works well for agentic work. Not for code harness. EDIT: ``` stronger on coding and raw capability but can be more argumentative, verbose, and costly. Reliability and instruction-following Many users say 4.6 felt more reliable and followed instructions better. "With 4.6, when I tell it something, it actually remembers the spirit of what I asked for and keeps applying it." Others report 4.8 drifts from preferences and can be frustrating to control. "I still find myself getting frustrated when it ignores preferences and drifts from instructions" Some people find 4.7/4.8 push back more and act more adversarial than 4.6. "The biggest complaint against 4.8 is that it is argumentative and "pushes back" constantly" Coding quality and capability Several users praise 4.8’s coding strength and thoroughness. "4.8 is technically impressive, especially for coding" Other reports say 4.6 could be better for certain coding workflows and breaks less. "4.6 still >> 4.8 for anyone else as well? Maybe I'm in the minority, but for my use cases Opus 4.6 is still better than" Some recommend mixing models: use 4.8 for key tasks and 4.6 for general work to save tokens. "What I do is... use 4.8 for key moments, and for everything else 4.6" Cost, speed and token behavior Users note 4.8 often uses more tokens and can feel slower because it “thinks” more. "4.8 is much more cautious, and as a result - slower. It checks everything, thinks for a long time etc." ``` [https://www.reddit.com/answers/601770d4-4059-478d-aa52-b445c...] | |||||||||||||||||||||||
| |||||||||||||||||||||||