| ▲ | manmal 2 hours ago | ||||||||||||||||||||||
My biggest learning after some experiments - a BF16 (unquantized) Qwen beats a Q8 of double its size for decisions. I guess that’s the reason Kev switched to 4B BF16, from the original 8B version. Isn’t it interesting that quantization seems to mess with decision accuracy? | |||||||||||||||||||||||
| ▲ | girvo 2 hours ago | parent [-] | ||||||||||||||||||||||
That’s fascinating, but not that surprising to me. We act like quantisation is free “Q8 is basically lossless” is often said in the local LLM community, but it really isn’t. The trade offs are worth it, personally, and the damage to coding ability seems low: decision model approaches are stricter though Super cool finding! | |||||||||||||||||||||||
| |||||||||||||||||||||||