| ▲ | sigbottle 2 hours ago | |
It's interesting though that Q4 seems to be enough, is there a reason that 4 bit floats are good enough for inference? | ||
| ▲ | nottorp an hour ago | parent | next [-] | |
Is Qwen 3.8 at Q4 good enough? I tried to run 3.5 27b Q4 on what local hardware i had (only 8 Gb) and i was very disappointed. 3.8 wouldn't have fit in my VRAM and i wasn't in the mood to leave it overnight at slow speeds so I didn't try. | ||
| ▲ | MaxikCZ an hour ago | parent | prev | next [-] | |
New models are trained with 8/4bit quantization in mind. Going from "native" 8 to 4 isnt as big of a step as going from 8 to 4 if native is full bf16. | ||
| ▲ | amelius an hour ago | parent | prev [-] | |
3 is the magic number, and 4 > 3. (seriously, nobody knows why any of this works; it's just a matter of trying) | ||