Remix.run Logo
throwaw12 3 hours ago

I have a theory about this, what if we all became dumber after 4 months of heavy AI usage?

I remember how I enjoyed agents between December and February, something started changing around March.

I thought models are getting dumber, but benchmarks were convincing opposite, initially I thought maybe they're quantizing models for day to day use, but Opus 4.8 and Opus 5 seems worse models than Opus 4.6

pickledish 3 hours ago | parent | next [-]

You are very much not alone, and I don't think we're all getting dumber -- I kept using opus 4.5 all through the nonsense that was 4.7, 4.8, and 5, and kept having a good time :)

I think we just need to decouple "doing better on benchmarks" and "actually more useful to me", since they've clearly diverged

serf 2 hours ago | parent [-]

doesn't that just mean that either a) you're using the wrong benchmark to judge or b) the benchmark that YOU need doesn't exist.

KronisLV 3 hours ago | parent | prev | next [-]

> something started changing around March.

The economics catching up with the providers in regards to how much compute they can burn per request and have it make sense for them financially?

A sort of model collapse where Opus 5 seems to love throwing out long paragraphs of text and it needs to be "fixed" by changing the output style and other patches.

I'm not sure, it might also catch up to Kimi K3 and GLM 5.3 and the models that I'm moving to from Anthropic.

throwaway219450 an hour ago | parent | prev | next [-]

Benchmarks test whether models can pass exams with a right answer or a green test case. I don't think the models are getting dumber, but they're definitely getting more incomprehensible to talk to. I've noticed this happening almost as a step change with the overuse of words and tics, and so has the broader community apparently. We haven't all been getting dumb at the same rate.

blehn 9 minutes ago | parent | prev | next [-]

occam's razor explanation: more tokens = more $.

models are incentivized by their makers to burn through as many tokens as they possibly can, so long as the customer doesn't cancel.

bombcar 38 minutes ago | parent | prev | next [-]

Benchmarks for agents are entirely pointless and obviously so; I'm not sure why they even exist.

manojlds 2 hours ago | parent | prev | next [-]

If you re getting dumber, you would feel like it's all good right? Why would you feel Opus 4.6 is better than Opus 5

Foobar8568 3 hours ago | parent | prev | next [-]

I could stand Opus 4.6-4.8, I was impressed by the initial fable model. Codex 5.6 sol xhigh feels like the initial release of fable. Qwen 3.8 27b feels like using haiku or sonnet (I quickly stopped trying them).

unclebucknasty an hour ago | parent | prev [-]

>I thought models are getting dumber, but benchmarks were convincing opposite

>Opus 4.8 and Opus 5 seems worse models than Opus 4.6

After all we've heard about benchmark cheating, I'm earnestly not sure which or whether benchmarks are reliable anymore. But, beyond the models, I wonder if changes to their harnesses and/or instructions dumb them down. I have noticed models change their behavior, even when using the same version/effort. Sometimes for better. Sometimes for worse.

And, I have noticed a model go from really good to struggling. On 4.8 things were going well for a good stretch, so I did not switch to 5 when it came out. Even after hearing complaints about 5, 4.8 was still going well. Then, suddenly over the last few days, 4.8 seems to have nosedived. It feels similar now to the complaints I hear about 5.

In my case, it suddenly started ignoring my design guide, and introducing new fonts etc. It would even use several different fonts and sizes, as well as different margins for similar elements within the same page. It abandoned classes and started inlining styles. It started feeling random and, even after it realized it needed to go back to the design guide, it just continued with more of the same.

There seems to be something that happens after new model releases in both quality and behavior of previous models. It may not be immediately, but eventually there is frequently some regression.

Bluestein 37 minutes ago | parent [-]

Maybe compute relocation?