| ▲ | nsagent an hour ago | |
See this recent paper: Quantization Degradation in Large Language Models: A Signal–Noise Perspective [1].
This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will. | ||
| ▲ | merbanan an hour ago | parent [-] | |
I created a pruned experts model of the q2 quant, while it gave good performance on limited hardware there was severe quality degradation. | ||