Remix.run Logo
cyanydeez 17 hours ago

I assume all major cloud AI are lowering their quants and measuring user retention.

SwellJoe 13 hours ago | parent [-]

That doesn't really track, though. Even tiny models, like Gemma 4, have coherent prose. I know Opus isn't being quantized that small. It seems like it must be some kind of...I dunno. Maybe over-fitting toward some user metric that doesn't track how readable its writing is?

It seems to still be good at code (though I haven't directly compared to earlier Opus versions lately), so it's not a general model collapse type problem, nor quantization errors. If quantization problems, I would expect it to break down on logic before prose, since a 4-bit quantized Gemma 4, even the small versions, have pretty good and, more importantly, coherent written English.

cyanydeez 33 minutes ago | parent [-]

Qwen3.8 quant 4 27b works coherently and takes up very little space.

If the cloud AIs arnt downsizing the majority of their customers, they will be out of business.

cyanydeez 31 minutes ago | parent [-]

Qwen3.8-flash-next can run in ~60gb, offloading 50gb to ssd, and is comparible.

The business has to downgrade to be sustainable, and open models prove it.