| ▲ | meowface 21 hours ago | ||||||||||||||||||||||
They have repeatedly said they do not ever intentionally reduce model quality and do not degrade in this way, and that a model version number is always the same. But, of course, OP is an empirical claim to the contrary, and I'd be curious to see if anyone (who's been capturing data over these timeframes) can replicate the same results and if Anthropic has any comment. | |||||||||||||||||||||||
| ▲ | mh- 21 hours ago | parent | next [-] | ||||||||||||||||||||||
Every official statement I've seen around this is careful to say that they "don't intentionally reduce model quality", which leaves plenty of room for "we adjusted some knobs and our evals show performance is materially the same". However, I also agree that I haven't seen any robust data from someone tracking it daily/weekly. The handful of sites purporting to do this aren't even running it enough times to hit stat sig. edit: someone linked one elsewhere in this thread called AI Stupid Level - they "run 7 trials instead of just 1". I don't blame them. Doing this in a statistically sound manner would cost a small fortune. | |||||||||||||||||||||||
| |||||||||||||||||||||||
| ▲ | cma 13 hours ago | parent | prev [-] | ||||||||||||||||||||||
See March 26, 2026 incident. Model not degraded, but harness changed to strip out past thinking tokens when a session went out of cache to save money and ease capacity constraints (affected API users too), resulting in bad degradation. | |||||||||||||||||||||||