| ▲ | bitexploder 3 hours ago | ||||||||||||||||||||||
I suspect they play with their quants and perform weight sensitive tensor/parameter tuning among other things to get serving faster and some of the time for some workloads it surfaces. I feel this has a high probability of being correct and an explanation for some of this. | |||||||||||||||||||||||
| ▲ | Rapzid 2 hours ago | parent [-] | ||||||||||||||||||||||
I refuse to believe they "play with their quants" once a model version is labelled and shipped. What does that even mean; could you explain it please? These models aren't just used through claude/codex, they are used through API access and it's quite expensive. Previous regressions were related to harness regression, and platform issues. Not some Nerf conspiracy 99% of the vibe bros believe in. Note: I know what quantization is so don't hold back. | |||||||||||||||||||||||
| |||||||||||||||||||||||