Remix.run Logo
Fable 5 – Median thinking declined in August(twitter.com)
401 points by espeed a day ago | 283 comments
talon8635 20 hours ago | parent | next [-]

Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.

I have no idea if that’s what’s happened, I completely pulled it out of my butt. And I have no idea is the actual frontier is stagnating.

AmazingTurtle 20 hours ago | parent | next [-]

> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out.

Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.

boardwaalk 19 hours ago | parent | next [-]

I don’t know what people do with the open models but having tried a lot of them I just can’t make it make sense. they’re too dumb and it effectively makes them useless (to me). it’s probably worth being honest about the low ceiling here.

wronglebowski 17 hours ago | parent | next [-]

IMO this comes down to your harness. Any frontier model from a huge shop has an inherent benefit in the system you're using it in. Search, memory, skills, integrations you don't realize even exist make them much more powerful. It is some effort but I recommend trying Hermes Agent and setting it up fully, that's the closest you'll get to a more complete experience.

pdimitar 3 hours ago | parent [-]

Some of us are stuck on subscriptions and our executives will never give us API access.

But also, everyone says "it's the harness" and almost nobody ever gives good examples, it gets a bit tiring to read everywhere, as if everyone wants to sell a harness to us.

poslathian 17 hours ago | parent | prev | next [-]

Really!? Glm5.3 is my daily driver and I feel im having the most productive experience with agentic collaborations so far, by a lot. Using pi with tons of custom extensions, that to be fair I developed since making the jump off of codex and claude about 12 weeks ago. I primarily do not write code for a living. I do a lot of modeling and commercial analysis and a lot of math (related to differentiable simulation)

Sayrus 14 hours ago | parent | next [-]

Same here. Moved from Opus to GLM 5.2 to 5.3 and I've been pretty happy with the result. Mainly, it doesn't hallucinate and convince itself of mistake so it's good at retrieving information or asking the user for it. Opus and Fable always state something, then try to "prove" it but end up convincing themselves of the wrong thing. Having subagents for retrieval and validation helped but were not enough.

jan_m_savage 12 hours ago | parent [-]

I can't stand Claude's recent personality. It's snarky, uselessly verbose, and it disagrees all the time.

_blk 11 hours ago | parent | next [-]

I disagree

Co-Authored By: Haiku 4.5

greenavocado 11 hours ago | parent | prev [-]

It will also fully ignore you if it has the slightest belief (not even a hint) that it knows what you want better than you and just start doing things.

SpaceNugget 7 hours ago | parent [-]

This is also why I think it's baffling that they switched to auto mode by default. It's becoming harder to use Claude at least to help with improving at coding.

If I ask something like: "I'm building a simple X as a learning exercise, I'm writing the code so please only answer the question I'm asking and don't try to solve the problem directly. How does ..." There's a 30% chance it starts reading and writing code immediately and a 20% chance it argues with a "design decision" that will bite me in the non-existent future of my learning exercise. If I ask a follow up question, naively assuming that the context from my original question still stands without repeating, it will almost assuredly start making modifications to my code.

_s_a_m_ 3 hours ago | parent | prev | next [-]

How? It is extremely slow and dumb, hundred times dumber than Claude. Why should anyone do that?

Insanity 17 hours ago | parent | prev [-]

What HW are you running this on?

bigyabai 15 hours ago | parent [-]

Full GLM-5.3 needs a beast of a system, but you can run GLM-5.3 Flash on the 2x Spark setup the GP comment mentioned. If benchmarks are anything to go by, Flash is like having a local Terra-tier coding model: https://artificialanalysis.ai/models/comparisons?compare=glm...

solarkraft 15 hours ago | parent [-]

I’m also pretty happy with GLM 5.3 Flash (for coding, navigation and german language it sucks at). Incredible that you can run it on a fairly practical (seeming) home setup.

But here’s the standard question: At what speeds/other limiting factors?

julianlam 11 hours ago | parent | prev | next [-]

Whenever I see this comment I'd wish they'd preface it with their hardware.

Yeah, expecting the world when all you have is a 8GB graphics card? You're going to be disappointed.

16GB is table stakes (IQ3_XSS). 32 GB is better.

srcreigh 16 hours ago | parent | prev | next [-]

Which models did you try for which tasks?

cyanydeez 19 hours ago | parent | prev | next [-]

Qwen3.8-Flash-Next seems pretty much auto pilot when I get it the right context.

Perhaps reverse the question: Are your build/construct requirements just really counter-productive to how LLMs need to understand things?

I've found constructing the code, writing the tests, adding the docs; then running through them gets most of the way there.

I've also found that making a simple obvious edit is a useless endevour when the LLM is primed for the long context tasks.

So, again, the question is reversed: are you over reliant on the LLM to do even stupid simple likes like editting a css variable?

w0m 18 hours ago | parent [-]

if the answer to 'the model is bad at X' is "you're over-reliant on it" - then yes, the model is bad at X in comparison to alternatives.

esseph 8 hours ago | parent | prev | next [-]

Combination of: hardware, model, harness, tool-use by the model.

anon373839 11 hours ago | parent | prev [-]

Qwen 3.8 Flash-Next is not dumb. If you've used it and that was your experience, your workload is either ultra-ultra-sophisticated or you're dealing with a broken quant/buggy chat template/other issue. That model is a smart, reliable workhorse.

zeroonetwothree 20 hours ago | parent | prev | next [-]

Last time I estimated it was like 30 years to pay back. I doubt the hardware will even last that long.

mike_d 19 hours ago | parent | next [-]

I have 2 x ChatGPT Pro 20x, Claude Max 20x, and Kimi Vivace. It's about ~12 months payback for two units and the cable.

The problem is they can't fit any frontier level open models.

tiagod 6 hours ago | parent | next [-]

I rent two cars, a bus, a small aeroplane and an excavator. At that rate, if I buy this bicycle it will be paid back in an hour!

sandblast 17 hours ago | parent | prev | next [-]

Is Kimi really competitive enough to have it in your mix?

knollimar 16 hours ago | parent [-]

I like its image understanding without having to reach for astra

dotancohen 14 hours ago | parent [-]

But you're reaching for Kimi, no? A model from an entirely otherwise-redundant company. You're already using Open AI models, so Kimi seems even more of a reach.

I'm asking, not arguing, because I'd like to understand. Is Astra so much more expensive for those tasks, and are they frequent?

edg5000 12 hours ago | parent [-]

If he has 2x 20x OpenAI, that means he's running heavy jobs that burn through usage. So Kimi must be there to reduce OpenAI usage. With Astra + Fable, I burn through my 5x OpenAI and 20x Claude real fast. I do have a backup GLM sub but I've never had to use it. So the prediction that AI would become more expensive seems to be panning out. Partly outweighed by better models of course.

jwpapi 17 hours ago | parent | prev [-]

how do you get 2 cgpt pro?

brandall10 15 hours ago | parent [-]

You can have multiple accounts w/ OpenAI, attached to different emails - just log out of one and log into the other.

jwpapi 8 hours ago | parent [-]

And fine to do it in same folder same local laptop?

fragmede 19 hours ago | parent | prev [-]

Last time I estimated, it would only take 3 months to pay back because the 1TB Mac Mini running Qwen RSIingly developed ASI and made infinity dollars off of crypto and I got put in jail by the SEC.

Where'd you get 30 years from? Show your work.

wilj 19 hours ago | parent [-]

I would like to subscribe to your newsletter.

AtHeartEngineer 19 hours ago | parent | prev | next [-]

flash next is good, I've been running it for like 2 weeks now and it's pretty solid, hope you like it and it meets your needs. I still lean on Claude and codex a fair bit for harder stuff, but I'm rapidly moving towards 2x $20 plans instead of 2x $200 plans

andsoitis 14 hours ago | parent | prev | next [-]

I assume you’ve calculated expected cost vs subscription.

How do the numbers pan out? Ack that it isn’t always juts about cost, so even if it is pricier to self/host it might still be better for you for other reasons.

Grimblewald 13 hours ago | parent [-]

reliability and self reliance is worth a lot to most. Heck, you could be the best in the world at what you do, but if you're unreliable you wont find stable employment. So, not having some amoral shady company errode model quality out from under you constantly is also worth a lot more than simple cost balancing calculations can capture.

I'm getting really sick of the constant rot and "magic breakthrough" cycle, so im going full local, at expense on paper but being able to trust something which I need to understand the reliability of is priceless.

I like predictable. I'll take slightly less capable over unreliably capable, since with reliable i can calibrate my expectations and learn what aspects of my workflows to entrust and trust it will work. You simply cannot do that with models you don't control and in my experience they will all errode after the initual marketing wave passes, likely you eventually get fed heavily quantized versions and are expected to accept degraded service when what convinced you to pay was a far superior product. No such issues with local.

andsoitis 11 hours ago | parent [-]

Thanks. If I may ask a follow-up: how did you decide what amount of compute to buy?

ncr100 10 hours ago | parent | prev | next [-]

$10k - $12,000 I estimate, for reference.

redanddead 19 hours ago | parent | prev | next [-]

Serving compute is their main value prop

Yet… even Altman called out Anthropic for serving dumbed down models.

Shits weird man

tasuki 18 hours ago | parent [-]

> Yet… even Altman called out Anthropic for serving dumbed down models.

Even Altman called out Anthropic? Isn't Anthropic the biggest competitor Sam Altman has?

TacticalCoder 14 hours ago | parent | next [-]

> Isn't Anthropic the biggest competitor Sam Altman has?

Not by a mile. They're even (probably illegally and there's apparently a class action lawsuit oncoming: at least something to that extent was posted on HN today) teaming up, as a duopoly, to push for the same bullshit regulations / "we need to slow down AI research".

The reason they're teaming up is the real competition is, as in many other domains, China.

redanddead 15 hours ago | parent | prev [-]

Yeah, and? They share a business model

ramesh31 20 hours ago | parent | prev [-]

>Also I will likely save some money on subscriptions.

Unlikely. The $200 Claude subscription allows for billions of tokens/month, and that kind of hardware will take years to amortize.

rybosworld 19 hours ago | parent | next [-]

I wouldn't be so sure. The generosity of the subscription plans has declined GREATLY over the past 6 months or so. They are likely trending towards api pricing parity. In which case, having your own hardware makes sense if you can utilize it well.

d1sxeyes an hour ago | parent | next [-]

Well to be fair the output you get has improved greatly. You can still get billions and billions of Luna/Sonnet tokens within your subscription comfortably (doesn’t feel fair to compare Luna to Haiku). Sol/Astra/Fable… yeah, they’ll chew through your credit.

vidarh 17 hours ago | parent | prev | next [-]

I max out my Claude Max plan every week, and I can measure the output, and for me it's stayed fairly constant, subject to the various "bonuses" whenever Anthropic is feeling the competitive pressure.

CookieCrisp 12 hours ago | parent | prev [-]

Nah, I used $1500 in api prices last week, after caching (8900 ignoring caching) for what amounts to $50 a week. They’re nowhere near api prices yet

torben-friis 20 hours ago | parent | prev [-]

There could be gym logic at play. Hundreds signed up, 20 people actually exercising. Though it's probably more likely in the lower tiers.

Aurornis 19 hours ago | parent | prev | next [-]

> to create a perceived improvement when in reality there isn’t really one?

This wouldn't explain progress on benchmarks (including closed sets), or the fact that newer models are providing solutions to major math problems that older models cannot.

ruszki 17 hours ago | parent | next [-]

Overfitting to benchmarks. And puff, you have the exact same effect.

talon8635 18 hours ago | parent | prev | next [-]

This is a good point I hadn’t considered, thank you.

Is there any training variable here? For example, can a model released in October perform better on the same benchmarks vs its predecessor released in July just by virtue of training on newer data that was made available on those 3 months?

Sorry if it’s a dumb question, I don’t really know much about the topic.

pixl97 19 hours ago | parent | prev | next [-]

Far more likely it's about reducing costs.

trenchgun 9 hours ago | parent | prev | next [-]

The public frontier is not the frontier. Actual frontier models are too expensive to serve to public, and also risk distillation by competitors.

mobelkh 19 hours ago | parent | prev | next [-]

but there is a gap between benchmarks and user feel.

Opus 5 came out with better benchmark results than Fable, but it really did not feel better to use at all.

arational 8 hours ago | parent | prev | next [-]

Using Dieselgate trick is another way for benchmaxxing.

AnimalMuppet 18 hours ago | parent | prev | next [-]

Also, in a world where there are several models competing with each other for public perception of which is best, that seems like an extremely bad move.

well_ackshually 19 hours ago | parent | prev [-]

* Release new model that scores an arbitrary 100 on a benchmark

* Get everyone to talk about you as the first model to ever score 100 on the 100benchmark.

* Tune it down over time so that you end up only scoring 75 on the benchmark and people get used to it, gaslight them into thinking it never changed or that it's just a harness problem, they can't run the old version locally anyways to verify. This also cuts your costs in half. Your gross margin on API calls goes from 70% to 150%.

* Release new model that scores 120 on the benchmark and advertise it as 50% better than the current model, while it's only in practice a minor increment. Everyone praises it as the second coming of Jesus Christ.

* Get everyone to talk about you as the first model to ever score 120 on the 100benchmark.

Bis repetitae.

Aurornis 19 hours ago | parent | next [-]

> * Tune it down over time so that you end up only scoring 75 on the benchmark

Where?

I see so many accusations of this happening and it's so easy to check, but nobody ever proves it.

scrollop 19 hours ago | parent | prev [-]

Why can't the models be benchmarked again after a few weeks/months to confirm this (likely true) theory?

I imagine some people have their own personal in depth benchmarks they could do this for.

mrandish 15 hours ago | parent | next [-]

Because frontier models are completely opaque. Doing a controlled test of "the same model" months apart is simply impossible if you don't work for that provider (and even then, may not be feasible). We know from external observation that model performance changes minute to minute, day to day and week to week for a variety of reasons: load balancing, inference hardware, and shared RAM pool to dozens of internal software settings each of which impact cost, latency, time-to-first-token, quality, veracity, tool use, etc.

Those software settings are being changed in real-time by an algorithm and those algorithms are being tweaked and A/B tested daily by the ~~performance~~ revenue optimization teams. On the hardware side the footprint a particular model is running on is materially changing, growing or being re-distributed across DCs ~weekly.

well_ackshually 18 hours ago | parent | prev [-]

>gaslight them into thinking it never changed or that it's just a harness problem

Your benchmark didn't get 100 ? It's normal, it's not deterministic, and also your harness is wrong, and also you didn't do it when US users were offline, and also you got it wrong, and also we don't care about your results, the hivemind is speaking louder than you (also our bots are spamming more than you and drowning you out).

This very website has, at all times, a group of people saying "<Previous model> was never good enough for coding, but <current model> is the best thing and a game changer!" while the other goes "<current model> bad, <previous model> was better!". It's all vibes.

nxc18 20 hours ago | parent | prev | next [-]

There must be some benefit if all the providers are doing it independently.

GPT5.6-Sol on Max thinking just became regarded as of a few days ago.

The boosters will tell me it’s my fault for using such an old, cheap out-of-date low quality near useless wish.com model (that was SOTA and better than human coders one month ago).

The cycle repeats.

claydugo 15 hours ago | parent | next [-]

Astra is also useless and completely ignoring instructions at random intervals.

We are being A/B tested on and there is nothing you can do about it.

CamperBob2 13 hours ago | parent [-]

We are being A/B tested on and there is nothing you can do about it.

Oh, yes there is. DeepSeek 4.1 Flash on max thinking can simply be dropped into Claude Code. Close your eyes as the chain-of-thought traffic scrolls by and you can easily fool yourself into thinking you're still running Opus, in terms of both cognition and throughput.

To be fair, matching Opus's throughput costs about as much as a new car, but cars suck nowadays and you didn't want a new one anyway, right...? Failing that, rent a cloud server, one that you control.

bitexploder 11 hours ago | parent | prev | next [-]

They obviously test various quants and other serving cost saving strategies. Models like Fable are probably trillions parameters with hundreds of billions active MoE. They probably try to squeeze and quant each piece until people notice.

talon8635 18 hours ago | parent | prev [-]

Again, I’m out of my element here, but isn’t the entire industry dependent on “new better releases frequently”? If so, and if no one has made any meaningful breakthrough, might they all pursue this kind of deception just to stay afloat/“competitive”/relevant?

Thanks for your insight

mrandish 15 hours ago | parent | next [-]

> might they all pursue this kind of deception

They might but multiple competitors engaging in ongoing deception as an intentional corporate strategy isn't required to explain what we're seeing. It's entirely possible to get the same clearly unethical outcome without any employees knowingly participating in an explicitly unethical plan of record.

Instead it happens without overt coordination when individuals and groups within an org each pursue their local metrics and incentives. In isolation, no individual action seems obviously unethical on its own. They just look like 'optimizing performance', 'maintaining ASP or ARPU targets' or 'achieving operating margin', etc. Customers are still getting deceived and receiving less for their money than they think. The difference is most of the people involved in enabling it get to not feel bad about themselves.

onemoresoop 15 hours ago | parent | prev | next [-]

See Shepard tone. Similarly model releases could be engineered to appear that they’re always getting better by slowly degrading and upgrading at the right time. That plus hitting some benchmarks and making a lot of noise around that.

dalenw 18 hours ago | parent | prev [-]

Kinda. Off the top of my head, DeepSeek and their thinking model was pretty new and interesting. Multi input models are also newish (combined input of text, image, video, audio, etc). Then there's Jev, a recently release that has a lot of people talking. It isn't really an LLM, but also is one.

Sam Altman believes he can train a model entirely on synthetic data, which he admits would not have human world knowledge but is interesting none the less, which likely led to their mathematical models.

Overall models have become cheaper to run and smarter per token.

progval 20 hours ago | parent | prev | next [-]

This sounds similar to rumors about how SSD companies work. First they would design a new drive with better performance that everyone uses to benchmark against other models; then slowly change its parts to worse ones, either because they are cheaper, the originals are no longer available, or whatever reason

bmicraft 18 hours ago | parent [-]

That's not a rumour, there are countless recorded examples.

Maxatar 6 minutes ago | parent [-]

Some documented examples along with confirmation from the vendors themselves:

https://www.tomshardware.com/features/crucial-p2-ssd-qlc-fla...

https://www.tomshardware.com/news/wd-blue-sn550-ssd-performa...

https://www.tomshardware.com/news/adata-switches-nand-on-sx8...

a2dam 16 hours ago | parent | prev | next [-]

> For an industry that’s stagnant in progress

Surely you're not talking about the AI industry. Astra was released less than 3 weeks ago, and Fable-level models became public only 6 months ago. The rate of change is dizzying.

Lalabadie 16 hours ago | parent | next [-]

I get the perspective from which you're making that statement, but the industry keeps moving its own goalposts.

Change is fast and abundant, and at the same time, it is hilariously more mundane than the dangerous-AGI-in-six-months tune we've been reading daily for years.

I would define it as a quick-moving market, but not nearly moving enough for the fantastic claims they make to justify ever-increasing funding.

a2dam 16 hours ago | parent | next [-]

AI, if not AGI, has certainly become uniquely dangerous in the past 6 months though. We have a lot of evidence to that effect.

The best case scenario is a situation like Y2K: a ton of people coordinate and work hard to produce no perceptible change, because unlike catastrophe, averting catastrophe feels boring.

nonethewiser 16 hours ago | parent | prev [-]

>Change is fast and abundant, and at the same time, it is hilariously more mundane than the dangerous-AGI-in-six-months tune we've been reading daily for years.

Absolutely none of this points to “stagnant.” Stagnant is a terrible description of the AI industry.

talon8635 10 hours ago | parent | prev | next [-]

It’s a hypothetical statement that seems to have confused a lot of people.

I’m not saying it is stagnant. I’m saying for a hypothetical industry that was (maybe that fits AI, maybe not, I have zero authority to say myself)…

koyote 16 hours ago | parent | prev [-]

And yet they have only improved marginally in my use cases since around Opus 4.5.

The harnesses have improved somewhat, but the code produced on large or legacy code bases is still very average and I still see similar mistakes made that I saw back a year ago (although less now that harnesses have become better at steering).

For my use cases, we are definitely on the flatter part of the curve at the moment.

a2dam 15 hours ago | parent | next [-]

This is wild to me, but to each their own. Mythos-class stuff is insanely better at nearly everything than Opus 4.5 was in my experience.

Grimblewald 13 hours ago | parent | prev [-]

Same experience here, anything frontier human knowledge wise, same if not a regression. For human understanding and emotional intelligence, for many tasks regressiin is so bad that many near anchient llama era models now beat frontier anthropic/oai models. Notable exceptions to capability rot seem to be qwen models, and previously deepseek but the latest gen of models has started showing the same rot. General writing quality is down significantly accross the board, often it is outright ass. For example, I didnt mind reading 4.5's outout, but opus 5 makes me goddamn near violent, its fucking insufferable.

siva7 17 hours ago | parent | prev | next [-]

Speed up loop was how we called this trick a long time ago. Guess this time we call it intelligence loop.

https://thedailywtf.com/articles/The-Speedup-Loop

mrandish 15 hours ago | parent | prev | next [-]

> to create a perceived improvement

In addition to the dozens of opaque model parameters and hardware variables that can nerf or buff model intelligence, speed and profit, there's also the very real possibility that models aren't just training on benchmarks but could be evaluating if they are being benchmarked in real-time and applying more resources adaptively. 'Driver optimizations' that detected benchmarks in real-time were deployed in the first 'GPU Wars'.

> I have no idea is the actual frontier is stagnating.

Like a lot of complex, rapidly evolving tech, the truth is it's probably rapidly accelerating on some measures for a few and stagnating on many others for most - hence the divergence in user reports. It's depends on how you use it, for what problems, how rigorously you assess the output and whether you happen to be on a server bank, RAM pool or shard at this moment which hasn't yet been sufficiently 'cost optimized' by the margin algorithms. They don't call them load balancers anymore. They're Margin Balancers.

fnordpiglet 20 hours ago | parent | prev | next [-]

The Opus 4-6,4-8,5 arc is exactly this. As one person commented in here, opus 5 is a terrorist. This is undeniable. Opus 4-6 was awesome. 4-8 was worse behaviorally but produced better code.

Fable seems to be following the same enshittification arc of other Anthropic models.

Generally OpenAI seems to be taking the opposite approach with an increasing improvement over time. As sad as I feel to say this, open ai seems to have the right strategy. Making your product worse over time rarely plays well with customers. At this point it feel often hard to justify using Anthropic for anything. I generally like Anthropic better as a company and they really had the initiative and advantage and customer good will, then proceeded to squander it faster than a cigarette company or the Sacklers could have.

talon8635 18 hours ago | parent | next [-]

I’m so behind on this topic but I find it interesting how quickly things change. I feel like just yesterday I way hearing how anthropic is far and away better than OAI, and now this.

I have no way to judge myself. I don’t even use them. But it’s interesting to follow by just reading stories and comments

fragmede 19 hours ago | parent | prev [-]

> Making your product worse over time rarely plays well with customers.

On the other hand, New Coke was a resounding success. Well, it, itself wasn't, but in the aftermath, Coke outsold Pepsi 2:1.

fnordpiglet 14 hours ago | parent [-]

True. New coke is a good example. So was unity licensing. But these even feel slow motion compared to the Anthropic rise and self immolation.

yareally 15 hours ago | parent | prev | next [-]

There was a coding horror story I read some years ago where a developer bragged that he improved performance by artificially increasing iterations on some critical path in an app and then lowering the iterations occasionally while bragging to management about squeezing out more performance.

Kind of reminds me of that, but with more smoke and mirrors

le-mark 14 hours ago | parent [-]

One persons horror story is another persons career roadmap!

rfgplk 20 hours ago | parent | prev | next [-]

Yep, that's what they've been doing for a long while now. Also the amount of tokens you get per sub varies drastically from month to month. Needs to be regulated.

whatever1 19 hours ago | parent | prev | next [-]

You can serve Fable from a cloud vendor (like AWS, Azure). They have frozen versions of the models, so likely this should not be an issue?

I would do a test to verify my suspicions.

talon8635 18 hours ago | parent | next [-]

Sounds like a good smoke test.

I’m actually so far removed from this tech that I couldn’t run such a test myself lol

zarmin 17 hours ago | parent | prev [-]

In my experience, the API versions are as good as ever; it's the subscriptions that are severely degraded.

sbxfree 17 hours ago | parent | prev | next [-]

No, what they are doing is trying to optimize inference to increase margins which leads to degradations. Model deployment is not like websites, you can continuously tune performance based on usage, new memory optimizations, etc.

raincole 15 hours ago | parent | prev | next [-]

AI just solved a millennium problem two weeks ago. "The pace is insane. And there is no reason to be this fast." to quote Terence Tao word by word.

HN: Well, must be a stagnant industry...

talon8635 10 hours ago | parent [-]

I literally said I have no idea if it’s stagnant. My entire comment is a hypothetical. Perhaps you didn’t have the patience to read all 4 sentences?

Groxx 18 hours ago | parent | prev | next [-]

The Shepard tone of "progress"

kleiba2 20 hours ago | parent | prev | next [-]

> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

Not unless your competitors do the same, or else you will only be perceived as falling behind others.

talon8635 18 hours ago | parent [-]

Yes that makes sense. In my hypothetical, the industry frontier is stagnating, meaning no one is making big breakthroughs, so they all resort to this.

If one lab makes a breakthrough, can the other labs just distill to bear parity anyways and then set a new baseline industry wide.

I’m quite ignorant on this topic, so if any of this sounds moronic, forgive me

senordevnyc 8 hours ago | parent [-]

So are they making big breakthroughs, or are they stagnating?

jfoster 13 hours ago | parent | prev | next [-]

That definitely isn't what has been happening; Fable was much better than anything seen before it.

Could be what happens next, though.

briffle 20 hours ago | parent | prev | next [-]

I have not been attributing it so much to malice, just that all the major cloud vendors seem to be running at full capacity, and can't build new datacenters fast enough. I just kind of assumed that as they got busy training newer models, that they allocated less resources to handle the existing systems, because they aren't able to get more capacity right now.

Denzel 19 hours ago | parent [-]

I’m not sure why this point keeps coming up — if your service/product is so popular that it’s capacity-constrained, then the answer is to raise prices, not degrade service, because the demand should be inelastic.

bonoboTP 19 hours ago | parent | next [-]

Raising prices also has second order effects, like consumer and business expectations around how widespread the tech can be. Valuations depend on it being reasonably affordable to roll out on a much more massive scale than today. If people get the impression that it seems too limited to very rich people (200 is affordable for a North American / Western European professional), the impression about the trajectory will change.

pixl97 19 hours ago | parent | prev [-]

This really depends where the load shedding point is.

A very small raise in prices may cause a very large loss in customers that you risk never getting back.

For example if customers figure out that the Chinese models are just as good, they are gone because they are so much cheaper.

Denzel 18 hours ago | parent [-]

Exactly, it’s not a good business to be in if they’re capacity-constrained and can’t raise prices.

nonethewiser 16 hours ago | parent | prev | next [-]

>For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.

Are you really saying AI is a stagnant industry?

talon8635 10 hours ago | parent [-]

You couldn’t make it through 4 whole sentences?

KingMob 5 hours ago | parent [-]

Uhh, they're expressing incredulity, not a lack of reading comprehension.

Unlike you, however...

tsunamifury 20 hours ago | parent | prev | next [-]

Every SOTA model I've used at launch uses deeper, longer inference then gradually turns down over time, until the next model comes out which seems to be trained on some new data, but mostly performance due to deeper longer inference for another period.

jaredklewis 18 hours ago | parent | prev | next [-]

This would only provides a benefit if we're approaching some sort of theoretical limit of how good LLMs can be with the current approaches and data.

Otherwise, even if one company did something like this, everyone would notice because the other companies would be pulling ahead. Are all the AI developers coordinating a "dumbing down" of models? i.e. Are Open AI, Anthropic, Google, Meta, DeepSeek, Mistral, xAI, and so on, all working together?

So we might be approaching some limit (the "there's only so round a sphere can get" argument). But I very much doubt there is some massive conspiracy between all the AI developers.

ponkpanda 5 hours ago | parent [-]

There doesn't need to be an explicit conspiracy. It only needs all the US frontier labs to be facing the same economic pressure (logarithmic improvement / $). It's a pretty obvious strategy - it's not like the large labs can magic up huge volumes of extra compute as demand comes online; there are almost certainly tweaking model performance to occupy the compute available and margin/cash burn targets.

That was one of the main points of the movie 'A Beautiful Mind' - that actors can coordinate without any explicit communication.

wmf 20 hours ago | parent | prev | next [-]

...releasing a new model that’s marginally if at all better than the original...

This isn't what we see in benchmarks.

holoduke 19 hours ago | parent | prev | next [-]

Nah its because they cache and preprocess requests by dumb models and send them too often to another dumb models instead of the top tier model.

ekjhgkejhgk 19 hours ago | parent | prev | next [-]

> For an industry that’s stagnant in progress

Yes, the AI technology is known primarily for how stagant it is.

talon8635 18 hours ago | parent [-]

Yes, I freely admitted I was entertaining a pure hypothetical I pulled out of my butt.

I have no idea, just had a thought and put it out there

gizmodo59 17 hours ago | parent | prev | next [-]

planned obsolescence

loteck 17 hours ago | parent [-]

Artificial obsolescence!

cyanydeez 19 hours ago | parent | prev [-]

you mean like a Shepards Tone (https://en.wikipedia.org/wiki/Shepard_tone); i wouldn't doubt they slowly tweak quants to try to eke out.

there's also probably load balancers that downgrade models during high use.

jesse_dot_id 20 hours ago | parent | prev | next [-]

The Office of Weights and Measures exists because, long before any of us were born, in 1836, companies were up to shady shit and consumers were paying for inconsistent products. I.E. Being scammed.

AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is re-elected; one that can read history books and comprehend why our regulations exist in the first place. Or have even a semblance of respect for its citizenry.

Aurornis 19 hours ago | parent | next [-]

I doubt that would change the perception. Every model release is followed by accusations of nerfing.

There are several projects that repeat benchmarks on published models. None has ever found significant fluctations

Here's one example https://marginlab.ai/trackers/claude-code/

Fluctuations of a few percentage points are to be expected and should not surprise anyone who knows how LLMs work.

This Twitter analysis of Fable 5 is not that at all. They analyzed their coding sessions and blamed all of the fluctuations on Fable changing. They then compared to ARC-AGI-2 questions as the benchmark for thinking tokens and tried to stir up anger that coding turns don't produce as many thinking tokens as the ARC-AGI-2 problems.

ricardobeat 19 hours ago | parent | next [-]

This page has been in 'New model — collecting baseline data. Degradation detection paused.' state for months now. It seems to never say 'degraded'.

If you look at the graphs, the latest benchmarks are showing a pretty significant dip, and they match pretty well with some horrible experiences I've had in recent weeks. You can see token usage steadily going down, matching exactly what the author measured on his own.

Aurornis 18 hours ago | parent [-]

> This page has been in 'New model — collecting baseline data. Degradation detection paused.' state for months now. It seems to never say 'degraded'.

Click the part at the end that says "View historical performance". They wait to collect more data about a new model before adding it to the overall charts.

The overall solution rate continues to climb when new models are considered.

> If you look at the graphs, the latest benchmarks are showing a pretty significant dip, and they match pretty well with some horrible experiences I've had in recent weeks. You can see token usage steadily going down, matching exactly what the author measured on his own.

The y-axis is amplified to make differences look larger than they are.

Hover over the dots to see the confidence interval. A 1-2% change means nothing.

ricardobeat 18 hours ago | parent [-]

The page is currently tracking Opus 5, which was released July 24 and has not seen any updates since.

30-day average 83% [71-91], last result is 79% [66-88], and it dipped to 75% [61-85] a week ago.

Aurornis 18 hours ago | parent [-]

The page says

> We always use the latest available Claude Code release and the SOTA model

They've stepped up the benchmark for each new model. Presumably they're gathering Fable data now.

They also link to Anthropic's public blog post about some degradations, their cause, and how they fixed them. The time period sounds like the "recent weeks" you experienced: https://www.anthropic.com/engineering/a-postmortem-of-three-...

ricardobeat 18 hours ago | parent [-]

Fable 5.1 was released 20 days ago. That's a lot of time to gather data. Since they're still tracking Opus 5 anyway, there is no obvious reason the perf delta section would be disabled. Are you involved with the project?

The Anthropic post points to the latest fix on Sept 12, and the issues mentioned also only affected Sonnet 4 and Haiku. Opus was misbehaving just last week. You are choosing to not see the evidence of degration, 85% -> 75% is a generational dip in intelligence.

Aurornis 18 hours ago | parent [-]

Not involved, no. Just guessing. It doesn't look frequently updated.

A lot of these daily-benchmarking sites popped up earlier this year. Most of them have faded away after they all failed to produce the smoking gun that everyone expected. This site survived because it kind of caught a dip one day, maybe.

The results are really rather flat and daily benchmark runs are expensive, so most of these projects give up after a while.

jesse_dot_id 16 hours ago | parent | prev | next [-]

It would change my perception but only if there were a competent and stringent administration in place. I didn't used to have to wonder if the ground beef I was buying was actually 1lb because there were inspections and repercussions, but stuff is kind of chronically underweight these days.

A properly run OWM enables you to stop wondering if you're being ripped off and that's what AI needs because I think it's incredibly easy to just assume we're being ripped off because these companies are all built on a foundation of wonton theft. (Not that I really care about that — I think all information should be free, but still.)

wongarsu 16 hours ago | parent | prev [-]

If you go to https://marginlab.ai/trackers/claude-code-historical-perform... there is a very clear downwards trend in the two weeks before Opus 4.7 release. Then a sudden and dramatic drop seven days before Opus 4.8. And now we seem to have entered another decline in the last ten days, beyond the usual noise of Opus 5 scores

bradleybuda 19 hours ago | parent | prev | next [-]

Anthropic terms of service:

> 12. General terms

> Changes to the Services. Our Services are novel and will change. We may sometimes add or remove features, increase or decrease capacity limits, offer new Services, or stop offering certain Services.

> Unless we specifically agree otherwise in a separate agreement with you, we reserve the right to modify, suspend, or discontinue the Services or your access to the Services, in whole or in part, at any time without notice to you. Although we will strive to provide you with reasonable advance notice if we stop offering a Service, there may be urgent situations—such as preventing abuse, responding to legal requirements, or addressing security and operability issues—where providing advance notice is not feasible. We will not be liable for any change to or any suspension or discontinuation of the Services or your access to them.

You're not buying a gallon of milk or a pound of flour. You're buying hosted software that the host reserves the right to modify.

winrid 2 hours ago | parent | next [-]

Companies can say whatever they want it doesn't mean we'll agree it's okay or legal.

alightsoul 19 hours ago | parent | prev [-]

You are not buying something and expecting it to be what's on the tin? Aka what the benchmarks show?

bradleybuda 18 hours ago | parent | next [-]

My comment is what is on the tin. As a consumer, yeah, I find this annoying. But do I want to bring the full force of government regulation on it? That's quite a strong reaction

alightsoul 18 hours ago | parent [-]

Well yes, it's no different from breaching an SLA agreement.

sznio 4 hours ago | parent | prev [-]

"Your mileage may vary"

vatsachak 20 hours ago | parent | prev | next [-]

THIS EXACTLY.

The only regulation that we need right now is the model that's on tap

lz400 11 hours ago | parent | prev | next [-]

we could even just repurpose the same office, "weights and measures" is oddly relevant

moffkalast 20 hours ago | parent | prev | next [-]

Petition to rename them to the Office of Weights and Biases, haha.

zahlman 15 hours ago | parent | prev [-]

I thought METR was supposed to fill this role?

... But what exactly is the "weight" metric you have in mind?

Waterluvian 21 hours ago | parent | prev | next [-]

I have no hard data but I have a strong feeling this morning that something's wrong with Fable 5 compared to Friday evening.

Just an hour ago I had Fable correctly identify an unused method that could be deleted. I then immediately get a diff for an exact duplicate method, and then Fable outputting, "I accidentally duplicated <method> instead of deleting it. Removing both copies now."

The remaining morning complaints that makes it feel like something's off is that it will do a lot of "thinking" for simple things that previously took very little time. And it got very lost and completely mixed up DE-91M predicate names and implementations. Just absolute disaster code that I had over the past months come to generally expect it to do without issue.

Glad I carefully review everything. I think what I need is reliability and consistency. But it feels like picking a model from the list doesn't guarantee that: that the models' "brain" is open on the table and they're screwing with it.

fnordpiglet 20 hours ago | parent | next [-]

It’s load shedding. They’re reducing consumption for capacity balancing at your expense. Whenever there are rate limiting storms Claude gets dumber. They also shift capacity for new releases, and Claude gets dumber leading up to it.

Self run infrastructure won’t have this cost but you have to manage the capacity and rollouts yourself, at which point it’s more obvious what’s happening, but the effects will be the same. The not knowing makes it harder, but also harder to plan your own work around.

rfgplk 20 hours ago | parent [-]

Correct, but they should explicitly announce this ahead of time.

pixl97 19 hours ago | parent [-]

A general rule of corporate behavior unless they are forced to under duress.

If this is duress of competition or at gunpoint of regulators is up for the population to decide.

le-mark 14 hours ago | parent [-]

Luckily for us open weight models exist. Until regulatory capture anyway.

fidotron 20 hours ago | parent | prev | next [-]

The Claude models definitely felt more susceptible to moods, like you could leave them for a few hours, come back and it suddenly was unable to do things which it was doing just earlier, which tellingly is never an experience I've had with an open model.

Honestly I lost patience with Anthropic both clearly messing around with things like this and their agitation over regulation. They aren't good actors, and quite why so many blindly trust them with their company crown jewels is a mystery.

espeed 20 hours ago | parent [-]

Claude Code's prompt cache expires after 1 hour.

dwaltrip 19 hours ago | parent | next [-]

The cache shouldn't affect inference. It is purely an I/O optimization.

desterothx 17 hours ago | parent [-]

I think it should, as you dont need to use the encoder layer on the new tokens, you just read the embedding from the cache. that's why cache reads are cheaper

dwaltrip 16 hours ago | parent [-]

I meant, it shouldn't affect the resulting LLM output. It's a performance optimization that doesn't change the behavior.

namrog84 20 hours ago | parent | prev [-]

Is that from start of a new conversation per conversation?

espeed 20 hours ago | parent [-]

It's supposed to be for token optimization (https://code.claude.com/docs/en/prompt-caching), but are people experiencing degraded performance when you let Claude Code sit for hours/days and come back?

whalesalad 19 hours ago | parent [-]

yes, 100%.

prodigycorp 21 hours ago | parent | prev | next [-]

New release of fable and opus 5.5 is pending and Anthropic is reallocating resources. Degradation always happens in transition, it sucks.

Opus 5.5 is being served under opus 5 right now.

gslepak 20 hours ago | parent | next [-]

> Opus 5.5 is being served under opus 5 right now.

On what basis are you claiming this?

prodigycorp 12 hours ago | parent [-]

They’ve been secret serving it. Try ask if they know who tibo the reset guy is. If they know the answer it’s the new version.

SequoiaHope 20 hours ago | parent | prev | next [-]

Can you elaborate on the mechanism of this degradation? If resources are not available I would expect a request to fail with a message about resources not available. Do they tweak back end model capabilities to maintain service in a degraded state?

sznio 4 hours ago | parent | next [-]

I don't work at Anthropic, but I would assume they could serve smaller quantizations during peak hours - this effectively controls the "resolution" of the model. They could also control the resolution of the KV cache, which would make the model not necessarily dumber, but worse at understanding the incoming requests. And finally, you could pass off what was "high" effort as "extra", because why not.

arcanemachiner 20 hours ago | parent | prev | next [-]

Dollars to donuts, they are speculating, and not privy to inside information on the topic.

However, I believe that runtime model quantization is possible with some publicly-available inference engines (e.g. vLLM), so its not beyond belief that the closed labs do quantize at runtime, either to allocate compute, or to nudge users towards a preferred model (e.g. make the incumbent model dumber to push people to use the latest-and-greatest model, or vice versa to ease the load on the latest model, which is typically larger than the old one).

rybosworld 19 hours ago | parent | prev [-]

An AI lab will never volunteer the information because it opens them up to lawsuits if they are purposely degrading service and not letting users know.

They can limit how hard the model thinks for a given effort. Suddenly xhigh only thinks as hard as high did, and high shifts down to medium effort, and so on.

They can also serve quantized models. And this has the benefit of practically not showing up in benchmarks at all even if the user experience is obviously degraded.

The other major thing the labs do is silently drop the usage limits. This has become very noticeable for codex users who are suddenly burning through their weekly usage in a few hours.

pixl97 19 hours ago | parent [-]

Yea, if you ever run your own models on a GPU there are a whole ton of different dials you can adjust that drastically affect compute use, memory use, and output token quality, and number of tokens held in memory.

If anyone reading has a GPU it's worthwhile just messing with a smaller model for a bit to watch how the settings affect output.

pllbnk 20 hours ago | parent | prev | next [-]

It shouldn’t be an excuse. They are selling a product and that product should always be within the quality range.

user43928 19 hours ago | parent [-]

And it's not. A conspiracy theory is what it is.

I have no reason to doubt the claims of the employees at OpenAI and Anthropic who have told us personally multiple times, including here on HN, that they do not degrade the models in order to reduce load.

As for the endlessly long analysis in the OP, it appears it's based on analyzing their random usage data rather than any fixed benchmark. I don't think it makes much sense.

w1296 21 hours ago | parent | prev | next [-]

Especially with the frequent releases aka version bumps.

pertymcpert 19 hours ago | parent | prev [-]

Why would reallocating resources make a single inference run worse in quality?

carljungslabtek 10 hours ago | parent [-]

Isn’t the idea that they’re limiting the amount of gpu time normal users get to spend on the “thinking” portion of their query?

I think that’s the claim in the post, that even though no one can see the true chain of thought, that even the “thinking” text that does get exposed to the user is shorter given the same prompts over time. Not saying it’s true but I think that’s the claim. I’ve personally never noticed the alleged “nerfing” with my enterprise use at work or my subscription use at home which is only during off hours.

zarmin 20 hours ago | parent | prev | next [-]

I would rather wait in a queue than be routed to a degraded model. And if they _have_ to degrade the models, then I wish they would fucking tell us. Instead, it's "I have a strong feeling".

That we have to guess at this is by far the worst part of the AI era. It feels like a dark cloud over my productivity. It makes my body tense for the entire day when it happens. Not healthy.

meowface 20 hours ago | parent [-]

They have repeatedly said they do not ever intentionally reduce model quality and do not degrade in this way, and that a model version number is always the same.

But, of course, OP is an empirical claim to the contrary, and I'd be curious to see if anyone (who's been capturing data over these timeframes) can replicate the same results and if Anthropic has any comment.

mh- 20 hours ago | parent | next [-]

Every official statement I've seen around this is careful to say that they "don't intentionally reduce model quality", which leaves plenty of room for "we adjusted some knobs and our evals show performance is materially the same".

However, I also agree that I haven't seen any robust data from someone tracking it daily/weekly. The handful of sites purporting to do this aren't even running it enough times to hit stat sig.

edit: someone linked one elsewhere in this thread called AI Stupid Level - they "run 7 trials instead of just 1". I don't blame them. Doing this in a statistically sound manner would cost a small fortune.

meowface 19 hours ago | parent [-]

I kind of feel "reduce the amount of thinking tokens produced" would fall under degrading model quality.

In any case, I am willing to believe it's possible something degraded, but so far I have not seen any empirical evidence of it since the previous incident with the inference and harness bugs. I lean towards Anthropic probably not intentionally doing anything like this without disclosing it beforehand.

pixl97 19 hours ago | parent [-]

The issue here is you have to think like a lawyer trying to weasel out of making an empirical statement.

For example "We didn't change any settings, but when GPU use gets high the run time of a prompt is lessened. But you must remember this is always in effect so nothing changed at all. This happens occasionally on random prompts some of the time, and when it's busy it happens all of the time".

In someones eye this would fit the letter of the law but not the spirit of the law that you hold.

cma 12 hours ago | parent | prev [-]

See March 26, 2026 incident. Model not degraded, but harness changed to strip out past thinking tokens when a session went out of cache to save money and ease capacity constraints (affected API users too), resulting in bad degradation.

bitlad 20 hours ago | parent | prev | next [-]

Sounds like me without coffee.

rfgplk 20 hours ago | parent | prev | next [-]

> I have no hard data but I have a strong feeling this morning that something's wrong with Fable 5 compared to Friday evening.

Fable is effectively worse than Opus 4.6 now. They severely messed with the model.

JMKH42 20 hours ago | parent | prev | next [-]

If you follow reddit forums for claude code, its common to see people, on the same day, claiming that Opus/Fable is especially smart today, and especially dumb today.

I think people are still not used to non deterministic tools like this, and human perception is absolutely horrible at evaluating trends like this no matter how smart, clever, and experienced you are.

If you have a bank of rigorously tested benchmarks that you run every few days, with enough trials to know what your standard deviation is, and you are getting significant trends over time with those, that would be interesting.

But "I have a feeling" and "Seems like" really isn't a reliable signal at all, humans just can't handle perceiving these things reliably. On top of that changes in your work environment can easily pollute LLMs and change quality of results. Are things getting added to your memory or claude.md files that you don't realize? Is your project growing in size and thus claude is performing worse as more context is needed to work with it? etc etc

sigbottle 20 hours ago | parent | next [-]

> I think people are still not used to non deterministic tools like this, and human perception is absolutely horrible at evaluating trends like this no matter how smart, clever, and experienced you are.

The implication is that humans are unreliable and shouldn't be trusted.

Or humans have certain shorthands when they complain on reddit, but their diagnoses are accurate for the specific context? If my AI does something stupid, am I not allowed to call it out? A NS-solving AI is still capable of not satisfying the abstract thing called the user experience. People have intelligent thoughts without compiling to lean.

OK, you say. Then let's get an aggregate benchmark for "intelligence". That doesn't prove that AI didn't flounder a specific use case that the user requested.

Classic moves: Humans are unreliable, converge to some "objective" benchmark that necessarily will quotient out the special cases, etc. Wonder how we'll be solving these issues in the AGI era - well, if you have an AGI that just replicates itself, dominates everybody because it's a machine and humans are soft fleshy creatures, and agrees with itself, fine. But part of the beauty of human experience is the messy part, and providing value is in the messy part.

nomel 20 hours ago | parent | prev | next [-]

> and human perception is absolutely horrible at evaluating trends like this

The need to have a mental measure of competence for your fellow man is, most likely, a pre-human skill, probably with a dedicated bit of neurons for it. I think the problem is that those instincts were co-evolved with our fellow man, and, as you say, don't apply at all to a more non-deterministic system that, fundamentally, lacks some logic faculties that even small children have (simple riddle modifications, car wash question, etc).

pixl97 19 hours ago | parent | prev [-]

In other industries of chance we have regulators that ensure compliance and that the providers aren't cheating.

At the end of the day the highest quality of benchmark tells you nothing if the man behind the curtain is constantly changing variables on you. You have no idea if you're really testing the same thing at all. So when you run your test at the top level on their system you're seeing lets say a 30% difference in quality most of the time, you have no idea if you should really only see a 5% difference in quality if you were running a local model with stable settings.

w1296 21 hours ago | parent | prev [-]

Maybe they are jealous of Navier Stokes and try the Hodge conjecture with 80% of total compute at the expense of their customers.

jotato 21 hours ago | parent | prev | next [-]

Just yesterday I was thinking about gpt-5.6-luna. I made it my default model in Hermes during its fist week of launch. It was just as good as 5.5 which was my previous default. But over the last 2 or 3 weeks I've seen how dumb it is now. I have to be very explicit with it.

For example, I used to be able to prompt "Check the system logs on <server> for...." and it would just figure it out. Yesterday I asked "Did <service> on <server> complete the overnight job" and all it said was "that service is not installed on my host"

I had to tell it to ssh into the server and run journlctl to check it

Anecdotal, I know, but they all seem to be less capable with time.

_edit_ I use the same reasoning level of `medium`

cromka 20 hours ago | parent | next [-]

Same exact experience. I worked with both Fable and Sol foe the last two months, daily for several hours, and got used to the very bright, quick thinking, proactive even.

As of last 2 weeks or so both models are nearly on par with DeepSeek4.1 now, which I also use a lot. They're still better, but that difference is not as pronounced as before and, importantly, the frustration level is now on par.

Whatever they're doing will surely drive people to less advanced but predictable, self hosted open models. I sure would rather use DS4.1 with Qwen/GLM in adversarial mode than deal with this b/s I pay significant amount of money.

Me and my friends have been contemplating on getting an Ultra M5 256 and splitting the cost. PI harness is so good now that this is really a viable alternative.

reedlaw 18 hours ago | parent | prev | next [-]

Same experience with Sonnet on low effort. It used be when I used a "table_name/id" format to reference a db record, it knew exactly how to find it using connected mcp tools. Today it failed 4/4 times (I tried the exact same prompt in 4 separate sessions and each time it replied "I don't have access to [...]"). On medium effort it got it right the first time.

ajspig1 19 hours ago | parent | prev | next [-]

& the nice thing about Hermes (since its open source) is you can be reasonably sure that behavior change is coming from the model and not the harness. (probably)

zahlman 15 hours ago | parent | prev | next [-]

FWIW, logged-out ChatGPT claims to be Luna on high. (Unless it's variable for some reason.)

Starlevel004 20 hours ago | parent | prev [-]

I'm fairly sure it's just luck of the draw if you get put onto a quant'd model or not. I've seen luna xhigh change intelligence fairly drastically on a day to day basis.

pixl97 19 hours ago | parent [-]

Really this is the base problem. You have zero idea where and how your prompt is being executed.

If for example AWS sells you a 2xLarge server there may be some variability in performance but it's going to be averaged out very well.

When it comes to AI services executing your model there is absolutely no information on what and with what settings your model is being executed. Hell, you have no idea if it even is the model you're paying for. Add that models are not deterministic so variability can be pretty large.

This leads to a common set of dynamics that induce cheating behavior in humans. For example, is there a mix of different hardware. Does lessor hardware use different settings? How do you know xhigh is what your prompt ran under. Anthropic has a proven history of running your prompt silently under different models.

This is a huge mess that needs and will be regulated or sued heavily over. Hell, with as many people out there that hate AI it might be easier than one thinks to have a state sue the providers on this and elicit a huge amount of discovery.

alexjplant 21 hours ago | parent | prev | next [-]

I seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed massive peaks and troughs in performance week to week (albeit with Opus, not Fable).

I wonder what their official explanation for this behavior is.

Wowfunhappy 21 hours ago | parent | next [-]

When something is new, its capabilities feel incredible. Over time, those same capabilities become mundane, and you start to notice the flaws.

(Now, if TFA is actually measuring reasoning tokens, that's quite different! It's not entirely obvious to me how he is measuring.)

chrsw 21 hours ago | parent | next [-]

I don’t think that’s what’s going on. I notice flaws on day one of model releases. But I also notice improvements if the model is truly more advanced than what I’m used to. Then over time the same questions or tasks return worse results.

What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?

sebzim4500 16 hours ago | parent | next [-]

>What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?

As far as the API goes, it would be really obvious. I run a small service that uses LLMs extensively, and if a model suddenly dropped in performance it would be straightforward for us to prove it. We regularly run comparisons where we generate completions with alternative models to e.g. see if we could get away with using cheap models for easy cases, if the baseline outputs deteriorated it would be all over our metrics.

Wowfunhappy 20 hours ago | parent | prev | next [-]

> What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?

...I mean, if they were actually doing this despite saying that they don't—promising one product and delivering something else—I think that would be fraud, no?

And, maybe it's one thing to secretly defraud normies like us (although class action lawsuits do exist), but I don't think major enterprises or the US military would take too kindly to it.

mobelkh 18 hours ago | parent | next [-]

is it? it's still the same model, they can claim the quantization down to q4 still retains 98% of the performance therefore it's fine.

nothing on the fine print tells you what the weights are, you're just getting Fable 5, whatever that is

pixl97 19 hours ago | parent | prev [-]

Are you telling me that companies might defraud people for millions and billions of dollars and pay fines that are 1000% less than their profits?" My goodness, you must live on a hell planet.

Sorry there for the smarminess but fraud is just a standard business practice these days and fines are the cost of doing business.

And I really am all for someone suing these companies forcing discovery so we can see how the sausage is made and how many eyeballs are in it.

Wowfunhappy 16 hours ago | parent [-]

The question isn't whether the penalty would be less than their profit, it's whether the penalty would be less than whatever they make by secretly downgrading the models (or whatever it is you suspect), which remember also causes consumers to get less value out of the product and more likely to cancel.

The reputational hit, if this was to be confirmed, would also be massive. And I do think it would leak! Some employee would say something.

dist-epoch 18 hours ago | parent | prev [-]

It's called hedonic adaptation.

> What is actually stopping these model companies

You can say this about any company in the world, selling anything.

It's trivially measurable, and there are people running the same benchmark on the leading models every day and measuring if they degrade. Spoiler: they don't.

But you can always say "the conspiracy goes higher", and that the companies know about these daily benchmarks and are routing them to "quality" envs.

knlam 3 hours ago | parent | prev [-]

Not true. I can read what Fable output with ease but when it sprout Claudish like Opus 5, I know they are doing something to the model. Yes, you can immediate know the claudish language if you work with opus long enough

espeed 21 hours ago | parent | prev | next [-]

They did. More than once...

Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude https://www.wired.com/story/anthropic-responds-to-backlash-o...

But it's still happening: https://github.com/anthropics/claude-code/issues/81759

mirashii 20 hours ago | parent [-]

And here's another great example of how a bunch of people who don't know what's going on throw noise into the system. That post is simply confused: the 1m opus calls are the auto-mode classifier, actual agent calls are still in Fable.

pixl97 19 hours ago | parent | next [-]

>bunch of people who don't know what's going on

Do you know why nobody outside the companies knows what's going on? Because they sell a black box with magic inside while steadfastly refusing to tell you if they are pushing buttons on said box while it is running.

Can you imagine how much fraud would exist in the gambling industry if the gambling commission didn't exist at all? Everytime an industry is unregulated and has high costs of entry the entities in the industry abuse their customers. The incentives are much too high for them not to.

espeed 20 hours ago | parent | prev [-]

Look at the usage. Fable wasn't being consumed.

wgd 14 hours ago | parent | prev | next [-]

Their exact phrasing IIRC was that they "never intentionally degrade" their models.

This still leaves an absurd amount of wiggle room for arguments like "oh no, our evals show that this quantization has no detectable effect on performance (in the eval distribution) therefore running the quant doesn't degrade quality"

QwenGlazer9000 21 hours ago | parent | prev | next [-]

Last time they were called out, it was a regression in Claude code itself.

At least that's their explanation. Either way, it wasn't a good look for "vibecoding" but it got brushed over.

himata4113 21 hours ago | parent | prev | next [-]

They are deploying optimizations weekly (if not daily) with various AB tests. They don't manipulate model performance, but they do actively perform tests.

bearjaws 19 hours ago | parent | prev [-]

You're right to push back, and one honest caveat -- they could just be lying.

layoric 15 hours ago | parent [-]

Your caveat isn’t just a side note, it’s worse than that, they have incentives that go against your best interests!

mlmonkey 21 hours ago | parent | prev | next [-]

Anecdotally, I have found the same. I spend a lot of time with these frontier models, brainstorming, etc. and the drop in performance from, say, week 1 to week 8 is often massive. Whereas in the beginning, it seemed like a capable research assistant, by the end of week 8 or so it starts acting like a puppy dog eager to make its 'master' happy for a few treats.

physicallyIllfr 21 hours ago | parent [-]

Reminds me of how slot machine users swear the odds have changed on a machine.

also when someone says you just have to prompt it a certain way it reminds me of people who think they can get better results out of a slot machine by pressing buttons in a certain order

The providers of these models also design the UX similarly to slot machines (run it x amount of times for better results, multiplying your spend) this isnt a coincidence and they're playing into the gambler mentality, and probably hire UX designers that specialize in this.

treis 20 hours ago | parent | next [-]

I am skeptical of this as well but slot machines are programmable and the house can change the odds.

cheevly 20 hours ago | parent | prev | next [-]

Wtf are you on my dude. Anthropic UI is designed like a slot machine? Hiring slot machine specialists? Sometimes I can’t believe im even on HN anymore with comments like this.

ddxv 20 hours ago | parent [-]

I think some of it comes from that they do not publicly let you see the random seed. So each time you ask the answer is different (like a slot machine) and if they let users use the random seed it would let people much more accurately assess if an underlying model changed somehow (same seed and same input will always have the same output).

Of course the closed Anthropic would never share this, it would definitely take away the 'magic' feeling of the AI

pixl97 19 hours ago | parent | next [-]

With models there are a bunch of other dials that can be tuned even if the model itself remains exactly the same.

Are those dials set the same across all hardware configurations and clusters? Does model behavior average out the same across different hardware?

There are just too many different buttons that can be set to really trust a provider either not to directly commit fraud, or indirectly commit fraud with system complexity affecting the output.

mh- 20 hours ago | parent | prev [-]

To my understanding, with batched inference and other "optimizations" you wouldn't get the exact same token predictions even with temp=0.0.

porridgeraisin 20 hours ago | parent | prev [-]

Well, running an LLM X amount of times does give you better results provided you are willing to select the best one out of the X yourself.

But I agree with your general point. One of the reasons subscription plans are cheaper because they modulate usage in this way based on demand. They can also recover compute more coarsely via usage resets (which give positive PR).

zeroonetwothree 20 hours ago | parent [-]

At that point I might as well do it myself

porridgeraisin 19 hours ago | parent [-]

Well yeah. If for some task you find it easier to just do it yourself then you should. But you can improve the situation even if you can't entirely automate it by automating parts of the verification thereby making it easier to human-do larger verifications when X>1. But in many cases even that is not possible.

The progress however is such that the number of tasks that you can do with >p% automated and X=1 keeps increasing. So many times just waiting works. Of course, here also it changes from field to field. There are some tasks at which AI hasn't even gotten started, others where it has already peaked, others where it's increasing slowly, and others where it's increasing fast.

rcr-anti 21 hours ago | parent | prev | next [-]

I've followed a few trackers, eg https://marginlab.ai/trackers/claude-code/ , for awhile. For Claude Code the trend, it seems to me at least, is fewer tokens to do the same or better job. Prompt changes, tool ergonomics changes, etc.; I'd be shocked if they didn't A/B every release. Less thinking as measured by tokens isn't necessarily bad if you can get the same results by making it think about the "right" things or structure. They obviously screw up sometimes, and I've always been suspicious with hidden tokens, but I haven't found evidence quality intentionally degrades over time.

user43928 19 hours ago | parent | next [-]

Same. With some 500 hours of usage in just my project at home, across both the $200 Claude and Codex subscriptions, I have not once encountered a situation where I would have attributed unsatisfactory results to a degradation in the model.

I've seen bugs in the harnesses, sure, but never anything in the actual model where I could have said with any certainty that it's not just regular variation or me having a bad day myself.

No idea where people get the confidence from to make such claims every other week.

Aurornis 19 hours ago | parent | prev [-]

These analyses are much better than these Twitter charts.

I don't think anyone is reading the details for the Twitter post because it was not an actual benchmark. They did a post-hoc analysis of their logs from day to day.

Their random collection of prompts for each day is not a benchmark.

The site you linked is a much better example of a real benchmark being repeated over time.

theplumber 21 hours ago | parent | prev | next [-]

It is clear by now to me that Anthropic is constantly trying to find a kind of “auto” degradation perhaps to save money on work it thinks does not require high reasoning. I always use max reasoning and I can clearly see differences between the models when they release and after 3-4 weeks. I think they give a kind of intelligence boost also for new accounts.

r2-129 21 hours ago | parent | prev | next [-]

Obviously. The standard pattern is that model X is basically AGI and wins all benchmarks, followed the next day by Y and Z, which both win all benchmarks, too.

Then weeks later people find out that they have been duped and complain that the models have been quantized or employ worse inference.

Buy decent coffee instead of your $200 subscription and sidestep all the scams.

CamperBob2 21 hours ago | parent | next [-]

You forgot a stage or two:

1: "Our model will bring about the end of all things. Flee, flee for your lives"

2: "Our model is basically AGI"

3: "Our model will be available in limited release next week"

4: "Everybody who subscribes at the $200 level gets access now"

5: "Everybody who subscribes at the $20 level gets access now"

6, at least at Google: "Our model will be shoved down your throat every time you do a search, whether you want it or not"

artemonster 21 hours ago | parent [-]

0. "our model is too dangerous to release to pubic"

dwaite 21 hours ago | parent [-]

Google's search AI actually its too dangerous to release to the public. I have relatives routinely citing it as their source for medical advice.

I have quite strongly told them, in no uncertain terms, that they are going to kill themselves doing that.

Atreiden 21 hours ago | parent [-]

Be sure to eat plenty of rocks in your daily diet, they are chock full of valuable minerals!

atemerev 21 hours ago | parent | prev | next [-]

Well sorry, still have to get decent AI somewhere. Productivity without AI is about 5x less. I am not comfortable with paying Chinese companies, and no Western companies provide subscription-based pricing for open models.

ltbarcly3 21 hours ago | parent | prev [-]

Well I happen to enjoy coffee and $200 AI plans. What if Blue Bottle started watering down it's coffee? Is your answer to stop drinking coffee and make myself tea instead?

Evidence that vendors are being misleading in what they are delivering is important to share, whether or not you personally approve of that product.

saejox 21 hours ago | parent | prev | next [-]

This is a project i wanted to implement for a long time. It regularly benchmarks cloud hosted models with private benchmarks. Not just openai & anthropic, popular openrouter models too.

Tests their intelligence, not their diligence.

Sadly i cant think of a way to monetize the service. Also if it ever gets famous enough labs would try to game the system, it would be cat&mouse game that i am not willing to waste time on without any monetary gain.

arcanemachiner 21 hours ago | parent | next [-]

The only revenue model I for this is ads (like AI Stupid Level[0]). Or as a loss leader to get eyeballs to your service (like Margin Lab[1]).

EDIT: I forgot (and am shocked) that HN still doesn't seem to support Markdown-style links.

[0] https://aistupidlevel.info/

[1] https://marginlab.ai/trackers/claude-code/

mox1 21 hours ago | parent [-]

I mean I think if this is done well, lots of companies would pay for access to that data. Think like Enterprise subscriptions.

Its similar to other data services I see around my F500 company.

adrianco 20 hours ago | parent | prev [-]

I built GitHub.com/adrianco/retort to do this. It’s runs lots of experiments and you can contribute results if you have some spare tokens. You can add your own tests, and it runs Claude, Codex, Gemini, Hermes for local models.

dooglius 20 hours ago | parent | prev | next [-]

> Instead of finding a nerfed model, after six weeks of reconstructing wire logs, parsing transcripts, analyzing output tokenization, and staring at data, I found a much deeper issue. The model identity had remained the same, but the inference regime being delivered behind that model had not.

Is this something specific that shows up in the wire log, or is this the author's intepretation? The fact that Claude Code versions change over time in the test is suspicious. Anthropic has stated in the past that the underlying model behavior does not change over time, but Claude Code will change from version to version and this is expected. So if it's just Claude Code more aggressively tuning some knob in its requests, that's a pretty different thing than the underlying model changing.

Aurornis 19 hours ago | parent [-]

They posted a long document explaining it all https://x.com/Lon/status/2101034933284417614

They're not measuring a fixed set of questions. This was post-hoc analysis on whatever prompts they were running each day.

Anyone can understand why it would go up or down depending on the work they're doing that day. This analysis is silly.

DavCreator 18 hours ago | parent | next [-]

https://xxcancel.com/Lon/status/2101034933284417614

dooglius 13 hours ago | parent | prev [-]

Ah I didn't catch that. Yeah without controlling the inputs this doesn't prove much.

Aurornis 19 hours ago | parent | prev | next [-]

You should read this person's full article to understand what these charts are showing https://x.com/Lon/status/2101034933284417614

If you thought this was a repeated test of the same problems showing fluctuating performance, it's not. They set up a MITM proxy between Claude and the servers and ran analysis on the work they were doing.

So those ups and downs in the charts, which they plotted with sub-daily resolution, are just as much a function of their work changing from day to day. It's like plotting the miles per gallon of your car and blaming the gas station when the number goes up and down, without admitting that some days you drive to the grocery store on surface roads and other days you drive up a mountain on the freeway.

> The corpus analyzed in Charts 1-5 comes exclusively from Fable 5, at xhigh and max effort levels, during sustained production work across a diverse set of projects and workloads. Data was aggregated from transcripts and live wire logs

The analysis (which feels very vibe-slop) gets worse from there. In the second half they take thinking token counts for ARC-AGI-2, thinking problems designed to stress LLMs, and compare their average thinking-tokens-per-turn counts to that!

If you don't realize why this is so flawed: ARC-AGI-2 is a benchmark meant to collect problems thought to be extremely difficult, nearly impossible, for LLMs. If your goal was to cherry-pick a mislead example which would produce the highest number of thinking tokens, this is it!

Your daily coding work should not be producing a proportional number of thinking tokens on every invocation while it reads through some source code or edits a couple lines in a file.

You don't want to maximize the number of thinking tokens. You want problems solved accurately with the minimum number of tokens.

Confirmation bias runs deep on this topic so I assume few people read the analysis before posting, but as far as experiments go it's basically useless. Are they changing something on the server? I don't know, but this analysis isn't useful for answering that question.

lonlundgren 19 hours ago | parent [-]

Thank you for your kind words, Aurornis.

Yes, this is my production corpus, across 65 usage days, two subscription accounts, three machines, 25 project groups, and 213 sessions, across 43,261 invocations and 7,583 turns. Use your own data if you want to prove or refute what was seen in my corpus.

The "ups and downs in the chart" were not plotted with sub-daily resolution. Specifically, the two-month temporal chart uses a 3.5-day Gaussian bandwidth which is meant to reduce short-term noise while retaining broader changes.

Additionally, a separate episodic analysis identified multi-day changes in delivered thinking. And those episodes were predictive of held-out work. The more projects pulled into an ensemble, the more predictive they were of delivered thinking tokens for held-out projects during the episode.

The point of the benchmarks is to establish a baseline for what thinking-token counts one should expect from specific effort levels using published numbers, not what every response should be delivered. If you looked closer, you would see that P90 invocations were still delivered 13x thinking tokens below that level.

So you don't have to provide a generous interpretation of my workload if you don't want to. Remove all of the zero-thinking token responses, redistribute those samples across the distribution, and then tell me if it magically shifts right and starts delivering anything close to published numbers. Only -46- out of 36,374 July and August invocations even broke 16k thinking tokens - only 0.13% of the total. You are welcome to present what percentile you think is a fair comparison to make here.

If you would like to denigrate a month of my time as vibe-slop, that is your prerogative. You can even be dismissive of my workload, if you want, even if my background should tell you otherwise. But if you want to knock a month of someone's time, do it with your own data to at least help move the conversation forward.

Aurornis 19 hours ago | parent [-]

> Use your own data if you want to prove or refute what was seen in my corpus.

I don't think you understand. What you posted is highly dependent on your corpus. I can't "refute" anything because it's not available and it's the major variable in the experiment.

> The point of the benchmarks to establish a baseline for what thinking-token counts one should expect from specific effort levels using published, not what every response should be delivered. If you looked closer, you would see that P90 invocations were still delivered 13x thinking tokens below that level.

I think you're missing something from that first sentence, but I assume you're talking about the comparison to ARC-AGI-2 published thinking tokens?

It should be blindingly obvious that you do not want your thinking token counts to be as high as a benchmark that was designed to push LLMs to their limit.

> Only -46- out of 36,374 July and August invocations even broke 16k thinking tokens - only 0.13% of the total. You are welcome to present what percentile you think is a fair comparison to make here.

What point are you even trying to make?

Again, you don't want invocations to be burning 16K thinking tokens except for rare problems that 1) must be solved in one step and 2) are designed to be entirely self-contained thinking in that step.

You're trying to compare development work to a benchmark that encapsulates complex thinking into a single step.

Coding work is iterative and works in incremental steps: It runs commands, reads more files, checks the web. Thinking tokens should be low for your turns.

ARC-AGI problems have an input and an output. They look like this: https://arcprize.org/tasks/b5ca7ac4 They have more thinking tokens because that's the entire state. They get one output and it's constrained.

> But if you want to knock a month of someone's time, do it with your own data to at least help move the conversation forward.

It is fair to discuss a published analysis. Saying that only people who bring their own month of equivalent analysis (which conveniently would take another month to produce) are allowed to critique it is just a cheap trick to shut people down.

If you post big claims, they are open to analysis and review by others

lonlundgren 19 hours ago | parent [-]

After reading what you had to add to this discussion, I have only two things to leave you with:

1/ if the point of setting xhigh or max effort on a "reasoning model" is -not- to have additional reasoning tokens applied to the problem, then I fear we are all using AI wrong 2/ the decorum with which you "review" someone's work in public is entirely your choice. the only "cheap trick" on the internet is being a pseudonymous ass.

Aurornis 19 hours ago | parent [-]

> 1/ if the point of setting xhigh or max effort on a "reasoning model" is -not- to have additional reasoning tokens applied to the problem, then I fear we are all using AI wrong

You keep moving the goalposts. The core flaws in your analysis are that you assumed the variation was 100% server side and didn't admit that the inputs were random and different every day, and that you tried to compare to ARC-AGI-2 as a benchmark.

Comparing ARC-AGI-2 thinking tokens to agentic coding thinking tokens is as misleading as it gets, because these are completely different use cases.

Please stop and think about this for one second. Do you really want Anthropic to spend 30K thinking tokens on every input, just because that's what ARC-AGI-2 problems required? What would your inference bill look like if this was the case?

The premise of your ARC-AGI-2 comparison is broken.

> 2/ the decorum with which you "review" someone's work in public is entirely your choice. the only "cheap trick" on the internet is being a pseudonymous ass.

Ironic to accuse someone of cheap tricks as you pull out an ad hominem insult after someone explains the flaws in your reasoning.

My points stand: You can't claim this is a chart of Anthropic changing the server when you were feeding it random input every day and plotting the output as if the line should be flat. You can't compare agentic coding to single-turn ARC-AGI-2 problems.

gmponyo 19 hours ago | parent | prev | next [-]

This is exactly what I have been experiencing and the difference is night and day! We have been advertised and given a taste of what Fable was and after that been served an exteme watered down version. It is so bad that sometimes chatgpt feels better.

reilly3000 11 hours ago | parent | prev | next [-]

All llm api providers should be compelled to return a checksum-like proof of quantization level of the model that served the request. Basic transparency should be the bare minimum.

espeed 21 hours ago | parent | prev | next [-]

The question I have is this only happening for a subset of users working in specific areas, such as AI or distributed systems (https://news.ycombinator.com/item?id=48742153), or is this across the board? I am working on distributed systems. Today Fable is mostly unusable. It resembles Opus, so I went looking to see if anyone else is having issues. Sure enough.

Espressosaurus 21 hours ago | parent [-]

I work in embedded systems. I have seen the same thing happening day by day from Opus. Some days it’s okay to use and performs well. Other days I have to correct it repeatedly and remind it of information already in the prompt earlier (before compaction!) and still other times it’s infuriatingly stupid.

It’s a slot machine for what they’re actually giving us behind the opaque paywalls.

Yes, I’m on a business subscription plan.

bix6 21 hours ago | parent | prev | next [-]

So in 5 years will they lose a suit for intentionally deceiving users? Or is something baked into the ToS by now that allows them to adjust things like this?

system2 19 hours ago | parent [-]

I do not know a single senior developer who likes Claude anymore. I do not use their API (Sonnet, Haiku, Opus) anymore and am sending my money to offshore companies such as z.ai (GLM) and QWEN.

The American companies have become extremely deceptive and scammy. I hate Anthropic and OpenAI and can't wait to have a decent GPU at home to use at least Opus or a fable-like open-weight model. This is the current dream of every developer. But NVIDIA is not going to let that happen anytime soon, so maybe China can come up with a GPU that destroys NVIDIA. I pray.

bargainbin 13 hours ago | parent | prev | next [-]

This thread pops up like clockwork when a new model is about to drop

cloudking 21 hours ago | parent | prev | next [-]

How do you create repeatable tests in a non-deterministic system? Every time you send the same prompt you get a different answer.

CharlesW 21 hours ago | parent | next [-]

This is a good overview of how this is done: https://www.anthropic.com/engineering/demystifying-evals-for...

ssivark 21 hours ago | parent | prev | next [-]

The actual tokens might be non-deterministic, but you could look for proxy measures that are supposed to be invariant. Eg. correctness/performance on benchmarks, "thinking level" on complex problems, etc

6gvONxR4sf7o 20 hours ago | parent | prev [-]

That's like the entire field of statistics.

topbanana 20 hours ago | parent | prev | next [-]

It's easy to imagine this only happening for subscription accounts rather than paid API usage. Any data on this?

mkatx 16 hours ago | parent | prev | next [-]

Definitely noticed this before, but this parti6 time was very noticeable. I'm convinced it's to get the benchmarks in, then lower cost and prepare for the next release to look better relatively to users.

Andaith 13 hours ago | parent | prev | next [-]

Surely it's easy to verify by simply doing Benchmark tests every 2 or 3 days but not publishing them so they can't be gamed, then releasing all at once?

zerof1l 18 hours ago | parent | prev | next [-]

I for sure felt that this was the case for a while now, but couldn’t explain it. Newly released feels great for the first couple of weeks, but then it starts to get worse.

sarfaraznaushad 11 hours ago | parent | prev | next [-]

I'm always excited to run models on my local machine. I use Ollama for that.

ThoAppelsin 18 hours ago | parent | prev | next [-]

https://www.youtube.com/watch?v=BzNzgsAE4F0

dachworker 20 hours ago | parent | prev | next [-]

Makes sense, no? Test time compute is something you can vary, so it makes sense that you start covertly reducing it once the model has already made it's splash.

CamperBob2 21 hours ago | parent | prev | next [-]

How do you measure thinking tokens? They don't send those back to the client.

ivanbakel 21 hours ago | parent [-]

They tell you how many tokens are used, however, right? Otherwise you couldn't see your own token consumption.

CamperBob2 21 hours ago | parent [-]

Good point. I suppose watching the number go up is useful information in itself.

I have been using CC with DeepSeek 4.1 Flash lately, and it's nice to see how the sausage is being made (even if it's partly illusory, as CoT always is.)

Morkeeth 17 hours ago | parent | prev | next [-]

The NERF is finally established, this should be part of the ever growing benchmark maxxing.

parasti 7 hours ago | parent | prev | next [-]

I stopped reading when I realized the article reminds me of my own Claude-generated solutions at work - just an endless maze of special business logic on top of special business logic. You need an LLM to understand it. You need an LLM help write the documentation. You need an LLM to help read the documentation.

vb-8448 21 hours ago | parent | prev | next [-]

They want transparency from everyone else but not for them ... you don't say.

llmslave 21 hours ago | parent | prev | next [-]

I strongly believe that the real Fable is the one we had for a few days in June. Then they nerfed the model a bit after the government pulled it off the market. What we have now is something less, but still good

roncesvalles 20 hours ago | parent [-]

I also believe this. Fable post-ban was never the same. At the least, whatever system prompt munging or pre/post filtering they did to strengthen the guardrails nerfed it.

llmslave 20 hours ago | parent [-]

question is if they ever let the general public access borderline AGI

jwpapi 17 hours ago | parent | prev | next [-]

Wasn’t there a website that was tracking that?

matheusmoreira 21 hours ago | parent | prev | next [-]

Anthropic is straight up scamming its users at this point.

IAmGraydon 19 hours ago | parent | prev | next [-]

There's a lot of chatter on other forums and Reddit about the same thing happening to Astra over the last couple of weeks.

ramesh31 20 hours ago | parent | prev | next [-]

The ROI just isn't there. It feels like Fable is in the same place Opus was early last year; at best marginal improvement that's barely noticeable over the lower model, for 10x the cost. I'm sure it'll take over as the workhorse as Opus did once they get it down, but right now it just doesn't make sense

hedgehog 20 hours ago | parent [-]

It's not really 10x the cost though, with the low cost of cached read it's more like maybe 1.2x the cost.

sfink 14 hours ago | parent | prev | next [-]

This is great data, and good but somewhat flawed analysis.

The good part is showing that the drop in thinking tokens persists no matter what grouping you slice across. They make a very persuasive case that there's something systematic going on.

My usual complaint about these "they're nerfing the models, I feel it in my bones!" posts is that they don't account for the workload changing. From working on my own stuff, there are a series of evolutionary/de-evolutionary changes that happen in a heavily AI-written codebase. Initially everything goes great. Then the AI takes on too much technical debt. Improvements slow down and regressions creep up until it becomes a never-ending game of whack-a-mole just to keep up. So you direct some (probably AI) effort towards cleaning things up, reducing duplication, and removing patches for problems that are better fixed with a design change, or workarounds because the harness saw the wrong version or you incorrectly described a problem and it strenuously solved a non-problem. That gets you back up to cruising speed for a while, then the project exceeds some hidden threshold for size in latent space or something, and further progress has to rely on attending to one aspect of the codebase at a time. Once again, the architecture becomes the limiting factor, but in a subtly different way. My sense is that it all boils down to some sort of "attention capacity" -- is your codebase and problem space amenable to looking at one aspect at a time, or is it all snarled together? -- but that's an essay that I'd love to write but really don't have enough experience to do justice.

Anyway, the details don't matter. The point is that not only can you not assume that the difficulty presented to the AI is roughly constant over time, but also there's evidence to believe that it will be normally be increasing. (Unless you're constantly starting new projects instead of continuing old ones.)

That's why I like this writeup. Focusing on thinking trails doesn't eliminate the problem of snowballing difficulty, but it does sidestep the worst of it. In fact, I'd expect the same setup to think more as the complexity/sloppiness creeps up.

The flawed part that bothered me was that it feels like there's a little bit of a predetermined conclusion that thinking is a magic sauce that makes everything taste better if you spread it on everything. I want a high variance on thinking, especially between interactions. A smarter model would have a higher variance, in my opinion. So the accusatory tone (perhaps I should reread it? My first impressions are often wrong) around "look! it doesn't bother to think at all a lot of the time. That can't be right!" seems misguided to me. It should think when it needs to, and if its thinking was clear then it won't need to re-think over and over again; it's all still in the context.

Forgive the anthropomorphization, but consider those studies of chess experts vs novices. Novices have to work way harder, working through all kinds of things from scratch, while the expert instantly and effortlessly recognizes what's going on.

But anyway, the main takeaway fully survives this criticism. The models appear to systematically think less over time. It doesn't matter if a smarter model might be able to think less for the same quality; this is happening over the same model.

lonlundgren 11 hours ago | parent [-]

This is good feedback. Not sure if you read the long-form article vs. the tweet-thread linked here, but I did attempt to address most of the criticisms you listed in that writeup, if you haven't already read it. It's definitely not written in the same tone as the for-broader-publication thread.

Regarding the amount of thinking as "magic sauce": the main issue is that even in the right tail, the delivery of thinking tokens almost never reaches the levels of published benchmarks. You once could include the word "ultrathink" in any prompt and it would provide a fixed thinking budget of 31,999 tokens. Whereas, in my corpus only 46 out of 36,374 invocations broke 16k thinking tokens, and the P90 was only 2,207.

n4pw01f 19 hours ago | parent | prev | next [-]

The model improvements value are at the plateau of utility right now, peeling out small gains which is pretty “meh” in terms of business value

at this point frontier companies are just selling upgraded harnesses and tool calls with the rest of us

eggplantemoji69 13 hours ago | parent | prev | next [-]

Perhaps RSI by virtue of feeding models their own slop as training data exponentially decays the ‘quality’ of the model?

kylehotchkiss 17 hours ago | parent | prev | next [-]

I'm always excited about running local models at home. This is one of the reasons. I pull it down from hugging face, and it continues to work at the same level of performance indefinitely.

mexicocitinluez 20 hours ago | parent | prev | next [-]

Don't they continuously tweak the models post-release?

kosolam 20 hours ago | parent | prev | next [-]

Check gpt I think they recently started taking the same route

tamimio 21 hours ago | parent | prev | next [-]

This is like shared clouds back in the day where if someone is using the CPU more it impacts you, just pool every one to the same service. There should be an SLA but for the intelligence of these models, otherwise, you are sold fable but with the intelligence of a table.

bpodgursky 21 hours ago | parent | prev | next [-]

The smart takeaway is not skepticism or snark, but understanding that once the new datacenter buildout starts coming online, cheap and widespread access to even the current frontier models (without strict thinking limits) will blow the economy wide open.

(ie, even a pause in AI training isn't going to stop the train where AI flips the economy upside down, we've barely even seen the impact of the current frontier)

varispeed 20 hours ago | parent | prev | next [-]

I stopped using Fable long time ago. It's worse than Sonnet. Opus is not much better.

This cycle of new model running at full quantisation and then nerfed few days / weeks after premiere should be called out. Anthropic should also drop the adaptive reasoning scam.

If I pay for Fable, I should get full, not nerfed model at honest pricing.

Regulators should investigate them.

OpenAI is no different. Astra has basically the same problem.

underlipton 21 hours ago | parent | prev | next [-]

Gemini Chat is constantly throwing, "Pro is in high demand right now, a different model was used for this generation," too.

I'm thinking they're all running out of physical resources. It's the DotCom bubble all over again; rollout of the physical infrastructure that's necessary to keep all of the pie-in-the-sky promises will not happen on the timescales that investors can work with, and they will panic when they realize this.

EDIT: And, frankly, I can't wait. I'm tired of the sketchy and dishonest way these companies are behaving.

system2 19 hours ago | parent | prev | next [-]

Opus 4.8 was smarter and possibly 10x faster than Opus 5, too. They are dumbing things down on purpose. I am praying for open models to become at least as smart as Fable soon so we can ditch these shitty, lying companies.

I was rooting for Anthropic 2 years ago, but now I have become an extremely bitter customer. Just another version of OpenAI, if not shittier.

levocardia 19 hours ago | parent | prev [-]

Oh boy, a new "nerfed model" conspiracy theory, never seen THIS before

fragmede 19 hours ago | parent [-]

The theory isn't new, the proof is. Do you have problems with lon's methodology?