| ▲ | Rapzid 2 hours ago | |||||||
I refuse to believe they "play with their quants" once a model version is labelled and shipped. What does that even mean; could you explain it please? These models aren't just used through claude/codex, they are used through API access and it's quite expensive. Previous regressions were related to harness regression, and platform issues. Not some Nerf conspiracy 99% of the vibe bros believe in. Note: I know what quantization is so don't hold back. | ||||||||
| ▲ | r_lee 2 hours ago | parent | next [-] | |||||||
I would guess that if they do use such methods, it'd be to handle peak loads that go beyond their compute capacity, while they run the models at full capability when there's excess capacity like before Anthropic signed the Colossus deal, the usage limits were insane and everyone was complaining, I wouldn't be surprised if they'd rather try to make inference faster that way than try to just limit people, at least for those on subscriptions | ||||||||
| ▲ | dannyw 2 hours ago | parent | prev [-] | |||||||
Inference isn’t flat 24x7, peak hours have more usage, but you buy/rent servers; not servers only for peak hours. At their scale, you’d have to be setting money on fire if you’re not doing dynamic inference optimisations based on load. API and consumer subscriptions are treated differently; all trackers measuring via API won’t notice this. | ||||||||
| ||||||||