| ▲ | esafak 2 hours ago |
| Has anyone calculated the effective intelligence of these quantized models? |
|
| ▲ | nsagent an hour ago | parent | next [-] |
| See this recent paper: Quantization Degradation in Large Language
Models: A Signal–Noise Perspective [1]. We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation
This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will.[1]: https://arxiv.org/abs/2608.08188 |
| |
| ▲ | merbanan an hour ago | parent [-] | | I created a pruned experts model of the q2 quant, while it gave good performance on limited hardware there was severe quality degradation. |
|
|
| ▲ | mkl 2 hours ago | parent | prev [-] |
| There's some info in the README, including: > Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM. https://github.com/Niko1221/Strata#which-model-should-i-pick |
| |
| ▲ | nicce an hour ago | parent | next [-] | | I wonder how this Coder compares to Qwen 3.8 27B. Can it be really better since they are competitive for same memory requirements? | |
| ▲ | nisarg2 an hour ago | parent | prev | next [-] | | 92% is halfway to 99% Holds up pretty well | |
| ▲ | javier2 2 hours ago | parent | prev [-] | | ok that is getting interesting! |
|