Remix.run Logo
▲ girvo 2 hours ago

That’s fascinating, but not that surprising to me. We act like quantisation is free “Q8 is basically lossless” is often said in the local LLM community, but it really isn’t. The trade offs are worth it, personally, and the damage to coding ability seems low: decision model approaches are stricter though

Super cool finding!

▲ByteAtATime an hour ago | parent [-]

Interesting - I wonder if it's because coding doesn't use the specific token probabilities, while decision models do

▲anewhnaccount2 7 minutes ago | parent | next [-]

Yes this is the reason. Transformations like quantisation preserve the rank of outcomes much better than probability mass.

▲RussianCow 32 minutes ago | parent | prev [-]

I think it also helps that speed and latency aren't as vital for coding and we can afford to let the model think for longer.