Remix.run Logo
▲ jug 3 hours ago

We also have Nerf Bench:

https://www.bridgebench.ai/nerf-bench

They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.

This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.

▲comboy 2 hours ago | parent | next [-]

But you are using API not the CLI right? I did not ever observe API degradation, only subscription stuff through their CLI.

▲user3939382 2 hours ago | parent | prev | next [-]

Anthropic A/Bs my weekly quota amount. So I have an automated prompt that runs at 3 AM with a transcription task, I measure input and output tokens, and weekly/5 hour quota before and after. The absolute token counts stay within 0.1% while in mode A it counts for 1% of my 5 hour quota and mode B 4% of my 5 hour quota.

▲jacquesm 32 minutes ago | parent | next [-]

How did pissing off your customers ever become a business model?

I can't imagine sticking with a supplier that plays games like that with me. Tokens are a pretty vague quantity to begin with (you don't control how many tokens a model puts out in response) and giving a couple of purposefully wrong responses will happily inflate your bill, but you don't care because eventually it worked. It's almost an ideal vehicle to scam people.

Imagine the power company being able to decide how much you consume and at which price point.

▲herval 21 minutes ago | parent [-]

> How did pissing off your customers ever become a business model?

Airlines, banks, health insurance…

▲apitman an hour ago | parent | prev | next [-]

Do Anthropic quotas give you precise remaining token counts or something? I have something similar set up for tracking my ChatGPT usage but it only gives percentages remaining, which is a pretty coarse metric.

▲ffsm8 13 minutes ago | parent [-]

Claude code supposedly has otel you can set via env. I haven't set it up, so I'm just repeating hearsay.. but it supposedly has everything relevant in it wrt token usage and cost

It's meant for their test env I think, so is not documented to my knowledge

▲braingravy an hour ago | parent | prev [-]

Pretty amazing to see enshitification happen live with a product still in development… Truly web 4.0

▲LimitExperience 26 minutes ago | parent [-]

[dead]

▲avazhi 33 minutes ago | parent | prev | next [-]

Nerfbench isn't helpful if it's 3 days old.

▲andriy_koval 2 hours ago | parent | prev | next [-]

Usage bench is also very useful! Thank you for doing this!

▲Grimblewald 3 hours ago | parent | prev | next [-]

I dunno, I never sense nerfs for local models, but consistently a few months after launch for corpo hosted models, seems odd my internal model for the capacity of a model drifts for anthropic models but not local ones. I've been using LLMs heavily even before ada/babbage/davinci days, and trust my internal calibration over baseless handwavey explanations for why im imagining things, especially when I have data that shows capacity regression on frontier models for tasks, e.g. one shot success at loss, 0 success in 15 attempts once nerf is sensed. Others publish their quantified capability regressions which are also more trust worthy than this kind of handwaving.

▲eulgro 2 hours ago | parent [-]

Your comment makes no sense. How and why would a local model be nerfed anyway...?

▲r_lee 2 hours ago | parent [-]

he's saying that he notices a difference between local (not nerfable) and hosted ones, so that it's not as likely to be just placebo

▲jacquesm 31 minutes ago | parent [-]

That and 'loss' may have been intended to be 'launch'.

▲Rapzid 3 hours ago | parent | prev | next [-]

The vibe bro science is this always happens on every release, every Tuesday, and twice on Sunday.

Of course it's almost entirely unsubstantiated BS.

▲fbrncci 3 hours ago | parent [-]

Well now it’s being substantiated!

▲Rapzid 2 hours ago | parent [-]

Or rather it's being.. Unsubstantiated. The Nerf conspiracy isn't that there have been a few harness and platform bugs leading to performance regressions, but that OpenAI/Anthropic have maliciously and unethically degraded their model performance post release to shed load and save money.

▲somenameforme an hour ago | parent | next [-]

Yeah it's just inconceivable that companies whose entire business model started by engaging in wholesale for-profit theft and abuse of intellectual property would ever be so unethical as to try to lower their costs, especially just prior to an IPO.

▲Rapzid an hour ago | parent [-]

Again, these things are constantly measured. They sell HEAPS through their API access to enterprise consumers that expect a model to not be nerfed after it's released. And you bet many of those enterprises, some spending many millions each month, are measuring this shit.

So this is a case of extraordinary claims requiring extraordinary evidence.

And even though it's super straight forward to collect the evidence, there seems to be zero substantiating the vibe bro conspiracy theories.

▲mrandish 8 minutes ago | parent | next [-]

> They sell HEAPS through their API access

The claim is that they nerf subscription accounts not API.

▲somenameforme 18 minutes ago | parent | prev [-]

If you're going to try to argue that companies doing things, completely legal mind you, to increase their profit margins is a conspiracy theory then you're not debating in good faith. Let alone when we're speaking of a subset of companies that were fundamentally built on wholesale unethical behavior carried out for profit. Let alone when we're speaking of companies who are all racing to IPO where short term results matter more than just about anything.

Another issue is also that the risk here is probably literally zero. Any evidence in support of such could easily be dismissed, with completely plausible deniability, as a short-lived technical glitch as opposed to intentional behavior.

▲Rapzid 6 minutes ago | parent [-]

> an explanation for an event or situation that claims a secret, powerful group is responsible for a hidden plot, rejecting the standard or official account

I'm sorry, but yeah. The official account is a harness regression and some platform bugs.

Where is the evidence they are underhandedly and unethically regressing their models to shed load and reduce costs? This is the conspiracy theory running rampant through the vibe boroughs; that they are bait-and-switching on model capabilities then "nerfing" them to save money and shed load. Where is the evidence?!

I'm not saying it's illegal, per say, so don't come at me with that straw man bull cock. This bro science conspiracy has been circulating for at least 2 years(I don't even know) and enterprises would certainly be pissed off if they were paying premium API prices for advertised and previously tested model capabilities that are suddenly under performing due to "nerfing" shenanigans.

So where is the evidence?!

▲stackghost an hour ago | parent | prev [-]

To me that sounds exactly on-brand for Big Tech in general and Sam Altman in particular.

▲jackmott42 an hour ago | parent [-]

No one here has suggested that the conspiracy theory is stupid (But I will, it is stupid), we are pointing out that the people actually measuring model performance have not found the nerfing before launch conspiracy to be true. in short, yall dumb, shut up.

▲stackghost 16 minutes ago | parent [-]

I think the real story is just how easily people believe that purported conspiracy theory. It speaks to how little trust there is in these AI companies, and in Big Tech in general, that this "conspiracy" theory is perfectly plausible to lots of people

> in short, yall dumb, shut up.

no u

▲Razengan 2 hours ago | parent | prev [-]

Theory (Conjecture? Hypothesis?): What we notice as "model nerfing" is the company diverting compute to training/running new unreleased models..

Remember that some people get access to the next flagship version long before us peasants do. I recall seeing the mention of "Astra" more than a month before it was officially announced

▲Centigonal 2 hours ago | parent | next [-]

wouldn't less compute result in slower inference, rather than worse performance?

▲latentsea 2 hours ago | parent | next [-]

They could potentially quantize the model and run it at lower quality taking less VRAM.

▲poizan42 24 minutes ago | parent | prev | next [-]

My guess is that they are dynamically changing the quality of the model to always keep the speed above some floor. So once it gets below that they switch to a worse quant or reduce reasoning level, or some combination of both.

▲btown 2 hours ago | parent | prev [-]

The more likely thing that would happen is that the provider begins silently interpreting (perhaps some) high effort-level requests as medium, etc., or having a classifier do this far more subtly. As such, the load on the cluster is less, and more resources can be devoted to training. Whether the frontier labs actually do this is purely conjecture at this point.

▲JohnBooty 30 minutes ago | parent | next [-]

I assume there's classification going on where a really basic "Hi how are you?" style request sent to a high-effort instance can be routed to a lower-level instance. This... is pretty much fine with me, assuming they do a good job of it.

I would also assume they use nebulous labels like "Medium Effort" or "High Effort" map to quantitative amounts of compute allocation... and that these amounts can be varied manually or automatically. Right?

I mean, there's a reason why they call it "High Effort" and not "Exactly 5 Minutes of GPU Time on Exactly 10 GPUs." They want to be able to move those sliders and tweak those knobs.

▲nightpool 2 hours ago | parent | prev [-]

why is that more likely?

▲zxilly an hour ago | parent [-]

Because they already did so. The model in Codex will get lower `juice` than API version.

▲jackmott42 an hour ago | parent | prev [-]

There is no nerfing, look at the data before coming up with a theory as to why the nerfing that isn't even happening is happening.

fuck