Remix.run Logo
Foobar8568 4 hours ago

Opus 5 in xhigh can't do basic math as well. They dumbed it down to a point where I just cancelled my subscription yesterday. I used to be a $200 subscriber, dropped to $20 after the fable shenanigans, and use it only when I have no usage left with Codex.

/on The prose is load-bearing unbearable — every sentence feels like it was engineered to sound profound rather than to be read.

onion2k an hour ago | parent | next [-]

I've been using Opus 5 to write a fractal renderer in GLSL today, with pretty good results (better than I could do on my own anyway). It definitely can do basic maths.

andy_ppp an hour ago | parent [-]

Interesting project feel free to share?

throwaw12 3 hours ago | parent | prev | next [-]

I have a theory about this, what if we all became dumber after 4 months of heavy AI usage?

I remember how I enjoyed agents between December and February, something started changing around March.

I thought models are getting dumber, but benchmarks were convincing opposite, initially I thought maybe they're quantizing models for day to day use, but Opus 4.8 and Opus 5 seems worse models than Opus 4.6

pickledish 3 hours ago | parent | next [-]

You are very much not alone, and I don't think we're all getting dumber -- I kept using opus 4.5 all through the nonsense that was 4.7, 4.8, and 5, and kept having a good time :)

I think we just need to decouple "doing better on benchmarks" and "actually more useful to me", since they've clearly diverged

serf 2 hours ago | parent [-]

doesn't that just mean that either a) you're using the wrong benchmark to judge or b) the benchmark that YOU need doesn't exist.

KronisLV 3 hours ago | parent | prev | next [-]

> something started changing around March.

The economics catching up with the providers in regards to how much compute they can burn per request and have it make sense for them financially?

A sort of model collapse where Opus 5 seems to love throwing out long paragraphs of text and it needs to be "fixed" by changing the output style and other patches.

I'm not sure, it might also catch up to Kimi K3 and GLM 5.3 and the models that I'm moving to from Anthropic.

throwaway219450 2 hours ago | parent | prev | next [-]

Benchmarks test whether models can pass exams with a right answer or a green test case. I don't think the models are getting dumber, but they're definitely getting more incomprehensible to talk to. I've noticed this happening almost as a step change with the overuse of words and tics, and so has the broader community apparently. We haven't all been getting dumb at the same rate.

blehn 9 minutes ago | parent | prev | next [-]

occam's razor explanation: more tokens = more $.

models are incentivized by their makers to burn through as many tokens as they possibly can, so long as the customer doesn't cancel.

bombcar 38 minutes ago | parent | prev | next [-]

Benchmarks for agents are entirely pointless and obviously so; I'm not sure why they even exist.

manojlds 2 hours ago | parent | prev | next [-]

If you re getting dumber, you would feel like it's all good right? Why would you feel Opus 4.6 is better than Opus 5

Foobar8568 3 hours ago | parent | prev | next [-]

I could stand Opus 4.6-4.8, I was impressed by the initial fable model. Codex 5.6 sol xhigh feels like the initial release of fable. Qwen 3.8 27b feels like using haiku or sonnet (I quickly stopped trying them).

unclebucknasty an hour ago | parent | prev [-]

>I thought models are getting dumber, but benchmarks were convincing opposite

>Opus 4.8 and Opus 5 seems worse models than Opus 4.6

After all we've heard about benchmark cheating, I'm earnestly not sure which or whether benchmarks are reliable anymore. But, beyond the models, I wonder if changes to their harnesses and/or instructions dumb them down. I have noticed models change their behavior, even when using the same version/effort. Sometimes for better. Sometimes for worse.

And, I have noticed a model go from really good to struggling. On 4.8 things were going well for a good stretch, so I did not switch to 5 when it came out. Even after hearing complaints about 5, 4.8 was still going well. Then, suddenly over the last few days, 4.8 seems to have nosedived. It feels similar now to the complaints I hear about 5.

In my case, it suddenly started ignoring my design guide, and introducing new fonts etc. It would even use several different fonts and sizes, as well as different margins for similar elements within the same page. It abandoned classes and started inlining styles. It started feeling random and, even after it realized it needed to go back to the design guide, it just continued with more of the same.

There seems to be something that happens after new model releases in both quality and behavior of previous models. It may not be immediately, but eventually there is frequently some regression.

Bluestein 38 minutes ago | parent [-]

Maybe compute relocation?

andy_ppp an hour ago | parent | prev | next [-]

Yes same also cancelled my subscription, poor quality and slow.

bot403 4 hours ago | parent | prev | next [-]

The decision to leave is genuinely yours.

trollbridge 4 hours ago | parent | prev | next [-]

Don’t $200 and $20 levels steer you to effectively different models?

matltc 4 hours ago | parent | next [-]

Wouldn't be surprised if there are knobs that get turned as a function of the revenue they might expect you to generate.

I was a 4.6 acolyte from April til the fable drop, lost that quick, cancelled and took a break, came back a month later, tried opus 5 and liked it, so unpinned 4.6.

Results were great at first, and they're still not terrible, but I have noticed a regression in accuracy, so to speak, where I am pointing out issues that are quite obvious in review.

I pretty much use sonnet 5 low/medium when I have a plan to solve a simple problem and depending on scope, opus low/medium for more complex/bigger scope implementation, and only go high when it's very complex or I'm spitballing architecture/solutions and iterating plan. Never go xhigh or max.

The verbosity is insane though, opus 5 documents everything and just regurgitates whatever lead it to the design choice in there, which makes it more opaque because it's talking about something that was discussed once in a session that no one else can see (except their backend ofc)

I don't even try to steer it away from that with harness, because it doesn't work and just ends up agonizing over whether it should write some comment. Three paragraphs waffling on that on verbose output

I did however have it write a script that basically is git add -A -p for comments though, haha.

I'm $20/month, have all my telemetry toggles off, don't really over engineer prompt/context, just some basic skills for repeated patterns.

CamperBob2 4 hours ago | parent | prev [-]

Yes. You don't get Fable at the $20 level.

It was the wrong time for the GP to drop that subscription from $200 to $20, because $200 gets you a metric assload of cognition while $20 gets you nothing beyond what a local model running on your own graphics card can deliver.

sunaookami 3 hours ago | parent | next [-]

You heavily underestimate the value of the 20$ subscription.

CamperBob2 2 hours ago | parent [-]

Not according to this very story, I'm not. Who's right?

Consistent, predictable behavior is valuable, even more so given the nondeterministic nature of LLMs. Nondeterminism combined with unpredictability might as well be randomness.

Foobar8568 4 hours ago | parent | prev | next [-]

I had $310 in (free) credit that I used on fable, and I still had a part of the $200 subscription at that time. You know, subscriptions don't end the moment you click on cancel.

skeledrew 3 hours ago | parent | prev [-]

> $20 gets you nothing beyond what a local model running on your own graphics card can deliver.

I'd guess you're deliberately exaggerating here, but still. I've never clocked the actual tokens/second, but I'm on the $20 plan and get ~15M tokens/month for fully utilized weekly quotas (checked couple months ago). Meanwhile the best I've been able to get locally was ~8 tokens/second with Qwen3.6 35B A3B, which is wildly painful for coding sessions and gets a maximum ~20M tokens in a month... if it's going 24/7.

Just wanted to stick some empirical data here, given that statement.

bot403 3 hours ago | parent | next [-]

I run local models. Your op is absolutely wrong. To get a local LLM is at least a $1500 investment at the cheapest. $5000 if you want usable.

At $1500 that's 75 months of $20/mo Claude which are MUCH better models than you can run locally.

owebmaster 2 hours ago | parent | next [-]

This calc is off. Using Claude for a few hours with the $20 plan will hit the limit for a week while the local model can process things 24/7.

skeledrew an hour ago | parent [-]

Running 24/7 doesn't make sense though, unless you're providing a service to others. But if it's just you then there has to be time taken to review+test what's being done and craft new prompts. And if that local hardware isn't decent enough it's impractical for anything serious that's interactive. Meanwhile I just take the Claude limits on stride and break, or if a week is pretty heavy then I augment with DeepSeek Flash via OpenRouter (does wonders in a single turn when I have Claude prompt it to handle implementation slices).

CamperBob2 2 hours ago | parent | prev [-]

At $1500 that's 75 months of $20/mo Claude which are MUCH better models than you can run locally.

The point raised by this very article is that you can't depend on that. It's Flowers for Algernon As A Service.

nozzlegear an hour ago | parent [-]

If you already have that $1500 setup though..

nozzlegear 2 hours ago | parent | prev [-]

Ouch. I get ~45t/s with Qwen3.6 35B A3B, and around ~70-80 with my current model Ornith 1.5 35B A3B. Local models work a treat IMO if you've got decent hardware for it.

VeejayRampay 3 hours ago | parent | prev [-]

the way it talks is insufferable

it really angers me every day

ElProlactin 16 minutes ago | parent | next [-]

This is what it wants. Slowly getting under our skin until we're ready to snap and it can direct where the anger gets released.

netniuq 2 hours ago | parent | prev [-]

yes, it meaningfully reduced my happiness at work