| ▲ | guardiangod an hour ago | |
Extremely large 1 bit models are usually within 50-60% of KV divergence to lossless models. In this case I think the comparison to Opus 4.5 is a fair assessment. Extremely large models don't suffer as much from quantization due to its weight topology also contains encoded information, so the loss of info from any one weight is somewhat mitigated. | ||
| ▲ | ilc an hour ago | parent | next [-] | |
Any one weight, but all of them. And also crushing the architecture itself? I wouldn't pick up 400gb of hardware to run in that mode. I might try it for fun, but even then you are looking at handling a 95GB active parameter set. This is NOT a model for most home labs. I'm sure some can and will use it. But most, should steer clear. | ||
| ▲ | dist-epoch 11 minutes ago | parent | prev [-] | |
KL divergence (you misspelled it) doesn't tell you anything about capability drop - how much did this particular benchmark (thus ranking among models) change after 10% or 50% KL divergence? | ||