Remix.run Logo
▲ johnfn 3 hours ago

"Nerf"ing models isn't real in the vast majority of reported cases. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was.

I made a graphic to explain why people feel like the models get nerfed:

https://x.com/thesilenceturns/status/2103551351825543610

The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.

▲prodigycorp 3 hours ago | parent | next [-]

Incorrect.

Anthropic has admitted to nerfing in the past. There have also been inference bugs. On top of that, model performance changes as they move compute to schwaggier providers as well.

Your chart is wrong.

▲simonw 3 hours ago | parent | next [-]

> Anthropic has admitted to nerfing in the past

Where?

▲QwenGlazer9000 2 hours ago | parent | next [-]

Earlier in march/aprile, there was a regression in Claude code.

Unintentional tbf.

▲prodigycorp 3 hours ago | parent | prev [-]

Man this was last year and some Claude subreddit drama that I can’t furnish offf the top of my head but maybe one of the historians remember it.

▲p-e-w 2 hours ago | parent | next [-]

Noone can seem to remember anything with certainty when asked to actually substantiate these claims.

▲computerex 2 hours ago | parent | next [-]

There was this: https://www.reddit.com/r/Anthropic/comments/1sl5wfh/the_degr...

▲prodigycorp 2 hours ago | parent | prev | next [-]

I’m typing from my phone and im not going to review the semantics of Anthropic’s storied history of performance issues.

It’s not just ant. There are so many small knobs that providers can claim isn’t nerfing but “load management” or “improving user experience”. One example from OpenAI is reducing juice to reduce time to first token.

▲winwang an hour ago | parent [-]

You can just have your agent find the evidence, review it, copypaste it.

▲erinnh 2 hours ago | parent | prev [-]

I mean there is a direct link two comments down from here from 30 minutes before your comment: https://news.ycombinator.com/item?id=49902477

▲p-e-w an hour ago | parent [-]

That’s NOT Anthropic admitting to “nerfing” their model as claimed above (which implies intent), that’s a regression which they quickly fixed.

Christ this forum has become intellectually dishonest.

▲consumer451 3 hours ago | parent | prev [-]

I asked a historian:

Two postmortems, neither quite "admitted to nerfing":

Sept 2025, infra bugs: "A small percentage of Claude Sonnet 4 requests experienced degraded output quality" [0], alongside "We never reduce model quality due to demand, time of day, or server load." [1]

April 2026, Claude Code: default reasoning effort was lowered from high to medium, plus a caching bug and a verbosity prompt. Per Anthropic, "The models themselves didn't regress, and the Claude API was not affected." [2]

So users were right that quality dropped, but the confirmed causes were bugs and a product default, not deliberate model degradation.

[0] https://status.claude.com/incidents/72f99lh1cj2c

[1] https://anthropic.com/engineering/a-postmortem-of-three-rece...

[2] https://texxr.com/handle/claudedevs

source: https://claude.ai/share/4435bbcf-d6df-44a0-b1db-f08a11858bc2

▲what an hour ago | parent [-]

> bugs

There are no bugs, just happy little accidents.

▲consumer451 an hour ago | parent [-]

u/bcherny does sound a bit like Bob Ross now that you mention it.

▲johnfn 3 hours ago | parent | prev | next [-]

Sorry, you are correct - I modified my original post. I get frustrated every time there's a model release and 1 week later everyone is saying NERF! NERF! 99.9% of the time these people are wrong, but you are right that it's technically not 100% due to a few edge cases.

I am more skeptical about the compute provider claim - do you have any evidence of that?

▲r_lee 2 hours ago | parent [-]

I noticed that a few weeks back 5.6 sol would regularly glitch out and start speeding random words or loop and then the next day it'd be fine

and there's sometimes just huge floods of complaints from people all of a sudden, which is pretty unlikely to be a coincidence

▲swader999 3 hours ago | parent | prev [-]

Right, and it would be simple to un-nerf or shadow nerf by any kind of angle they want.

▲gobdovan 3 hours ago | parent | prev | next [-]

There are recorded cases of real regressions, but they're better characterised as incidents, not nerfs, e.g.: https://www.anthropic.com/engineering/april-23-postmortem

Btw, you have a typo in the twitter handle on your profile (not in your comment), 'thesilencesturns'.

▲johnfn 3 hours ago | parent [-]

You are correct, I softened the wording a bit. And thanks for the heads up on the typo!

▲winwang an hour ago | parent | prev | next [-]

That's a good observation, though I'd say here that two things could be true at the same time. But, I do personally believe that most of the reported nerfing is the case of your chart + latent evidence-less complaining. Honeymoon phases are real.

▲HawtAds 3 hours ago | parent | prev | next [-]

It's very much real but not necessarily malicious. We track upstream providers pretty closely. Sometimes it's a just matter of a single GPU runtime layer bug/update to break inference outputs. The model weights don't necessarily change/get quantized.

▲bitexploder 3 hours ago | parent [-]

I suspect they play with their quants and perform weight sensitive tensor/parameter tuning among other things to get serving faster and some of the time for some workloads it surfaces. I feel this has a high probability of being correct and an explanation for some of this.

▲Rapzid 2 hours ago | parent [-]

I refuse to believe they "play with their quants" once a model version is labelled and shipped. What does that even mean; could you explain it please? These models aren't just used through claude/codex, they are used through API access and it's quite expensive. Previous regressions were related to harness regression, and platform issues. Not some Nerf conspiracy 99% of the vibe bros believe in.

Note: I know what quantization is so don't hold back.

▲r_lee 2 hours ago | parent | next [-]

I would guess that if they do use such methods, it'd be to handle peak loads that go beyond their compute capacity, while they run the models at full capability when there's excess capacity

like before Anthropic signed the Colossus deal, the usage limits were insane and everyone was complaining, I wouldn't be surprised if they'd rather try to make inference faster that way than try to just limit people, at least for those on subscriptions

▲dannyw 2 hours ago | parent | prev [-]

Inference isn’t flat 24x7, peak hours have more usage, but you buy/rent servers; not servers only for peak hours.

At their scale, you’d have to be setting money on fire if you’re not doing dynamic inference optimisations based on load.

API and consumer subscriptions are treated differently; all trackers measuring via API won’t notice this.

▲Rapzid 2 hours ago | parent [-]

During peak hours requests queue and inference slows. During off-peak they can move systems over to training.

Where is the evidence they are "nerfing" the models due to request volume?

Edit: I don't know they do, I mean they could repurpose systems if they are idle. Inference demand is global, and providers like Azure have global routing options that are cheaper. Night time in the USA could be serving inference demand on the other side of the globe.

▲hbn 3 hours ago | parent | prev | next [-]

I have not been doing increasingly complex things since Opus 4.6 when models got really good.

My work at my job has stayed the same. But the model quality has varied.

They definitely tune the models in production after launch, if not only to share load during high traffic times. It’s not a crazy conspiracy that the same model can be stupider at different times.

▲frde_me 3 hours ago | parent | next [-]

> I have not been doing increasingly complex things since Opus 4.6 when models got really good.

This is a more a statement on the work you do and how you work versus the models. I'm doing more complex work since Fable (and now for way cheaper thanks to Opus 5.5)

With 4.6 I would still babysit a lot more code quality and so on. With the newer model I see myself talking about features at a higher level, and then not having to nitpick PRs to death. Which means most of my time is now spent talking to the model about the product instead of the implementation of the product.

▲usef- 2 hours ago | parent | prev | next [-]

What sorts of things, if you can say? Is it a similar sized/complexity codebase? Most projects do become larger and/or more complex over time. And most people's standards do creep up as they learn.

▲johnfn 3 hours ago | parent | prev | next [-]

It's not about doing more complex things - complexity is more dictated by how large your codebase is, etc.

> It’s not a crazy conspiracy that the same model can be stupider

Sorry, I really do think it's a conspiracy. If nerfing were real, it would be trivial to prove. DeepSWE, SWEBench, and other benchmarks are all available for anyone to run. A "nerfing" hypothesis has to survive the fact that a statistically significant dip in benchmarks has never been observed.

▲Computer0 3 hours ago | parent | prev [-]

Open AI admits to such here: Open AI aims to have a stable API and admits to meddling with effort levels and such for subscriptions -https://news.ycombinator.com/item?id=49804316#49809266

▲eek2121 3 hours ago | parent | prev | next [-]

Admittedly, I didn't click your link, however, based on what you've stated, there is some inaccuracy. All these big companies take your requests and the context, and route it based on the content, cost, etc.

What Anthropic presents as Opus 5.5 isn't actually a single model...it's Anthropic's ecosystem as a whole. If you are lucky, you get the top model handling your issues all the time, however, that never happens. What really happens is that your request and content are graded along with your subscription (example: API? subscription, if so, what tier? how much has the user used it? Do we trust the user? how much? how much are they paying? are they asking something we think is dangerous?) and your request and context are routed accordingly.

Anthropic isn't alone in this behavior, Open AI does it as well, just look at the respective subreddits on reddit for both if you need some examples, or just play around with the various models from both companies.

There are a few folks who've done some analysis on this (their findings were posted on reddit and X), and a bigger multi-national study is apparently coming, though I admittedly don't know their findings.

I guess the tl;dr is that Anthropic and Open AI are actually selling you "best-effort" routers, so you may or may not get the best in class model, and only they get to determine if you do or do not. No guarantees.

▲Grimblewald 3 hours ago | parent | prev | next [-]

Nerf is real, i think we initially get full precision models and later quants. My own logs show it clearly for opus 4.5 to 5, consistently a few months post launch, models start making quant based mistakes, like slipping in inappropriate tokens (e.g. chinese ones in english text) which doesnt happen at all in the first few months and regularly later. Additionally frontier problems previously done well start being done poorly, until later model variants where performance mostly holds, likely due to them training on your data reguardless of what boxes you tick.

My local models don't display that degradation, sensed or measured. They consistently perform equally to what I expect of them, precisely because they don't change.

How does twitter explain that? Is my internal model for expectation of capacity magically not drifting for local models but somehow is for anthropic api call based models?

▲486sx33 3 hours ago | parent | prev | next [-]

[dead]

▲physicallyIllfr 3 hours ago | parent | prev [-]

Its because they're addicts and addicts always grow numb and immune to their fix, needing more dopamine. They want to feel what it felt like the first time.

By the way, dont for a second think LLM hourly limits are all about revenue, they're playing into this psychology. They hire literal gambling UX designers, they want to turn you all into addicts. They want to make you reliant.

Want to run your llm like a slot machine? They'll let you do that spin the generation on a multiple, get 6x results, pick your favorite. Feel that high.

Just know you can get that same hit of dopamine by fostering your own intelligence and creating something with it. Token dealers are just selling you the shortcut, straight to the reward, short circuiting the the natural process.

Bad times ahead for many. This shit isnt good for your brain. And you all know the truth, you just wont admit it. Its doing damage, making you lazier, less intelligent.. Making you an addict.

▲mwigdahl 2 hours ago | parent | next [-]

That’s right, and it’s been like this ever since we stopped programming in assembly language. Programmers’ brains used to grow manly and strong on a strict diet of manual memory management and custom stack frame handling. Once we transitioned to soft, weak modern languages like C it’s been all downhill.

▲physicallyIllfr 2 hours ago | parent | next [-]

Ahh, the very original and compelling comparison of llms to compilers.

You're a genius, did you think of that yourself? Or did you local token dealer teach you that?

▲kdkdjcjejxowjdj an hour ago | parent [-]

Ahh, the even more original and compelling “no true Scotsman” argument.

A classic. You go, buddy, use all the cliches you want! Whatever you need to make you feel like a big clever H4cK3r.

▲physicallyIllfr 24 minutes ago | parent [-]

Wild that you signed up for a whole sock acct to make this comment lmfao

▲ 2 hours ago | parent | prev [-]
[deleted]
▲cheevly 2 hours ago | parent | prev | next [-]

You are on actual drugs my dude.

▲physicallyIllfr 2 hours ago | parent [-]

Keep frying your brain with llms.

▲cindyllm 29 minutes ago | parent | prev [-]

[dead]