| ▲ | ilc an hour ago | |||||||||||||
To compare a 1 bit quant to the full fat model is misleading. Honestly this model people at home can tinker with, if you have a big enough Mac. Maybe 4 Strix Halo/DGX Spark, and then at 1 bit quant? Nah. Use the right sized model, for your hardware. You'll get better results. | ||||||||||||||
| ▲ | guardiangod an hour ago | parent [-] | |||||||||||||
Extremely large 1 bit models are usually within 50-60% of KV divergence to lossless models. In this case I think the comparison to Opus 4.5 is a fair assessment. Extremely large models don't suffer as much from quantization due to its weight topology also contains encoded information, so the loss of info from any one weight is somewhat mitigated. | ||||||||||||||
| ||||||||||||||