Remix.run Logo
▲ esafak 2 hours ago

Has anyone calculated the effective intelligence of these quantized models?

▲nsagent an hour ago | parent | next [-]

See this recent paper: Quantization Degradation in Large Language Models: A Signal–Noise Perspective [1].

  We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation
This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will.

[1]: https://arxiv.org/abs/2608.08188

▲merbanan an hour ago | parent [-]

I created a pruned experts model of the q2 quant, while it gave good performance on limited hardware there was severe quality degradation.

▲mkl 2 hours ago | parent | prev [-]

There's some info in the README, including:

> Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.

https://github.com/Niko1221/Strata#which-model-should-i-pick

▲nicce an hour ago | parent | next [-]

I wonder how this Coder compares to Qwen 3.8 27B. Can it be really better since they are competitive for same memory requirements?

▲nisarg2 an hour ago | parent | prev | next [-]

92% is halfway to 99%

Holds up pretty well

▲javier2 2 hours ago | parent | prev [-]

ok that is getting interesting!