Remix.run Logo
pizza234 4 hours ago

Their mention of the 5090 is bit odd, since on 32 GB GPUs, Q6 fits while having better quality. Very interesting model for 16 GB GPUs though!

sisve 4 hours ago | parent [-]

They mention 5090 with regards to speed, Q6 will not have that speed?

And speed matters a lot for many use cases

selectodude 3 hours ago | parent [-]

150 tokens per second on a ternary model implies that it’s GPU bound, I’d bet a Q6 model is even faster because it’s existed longer and seen more optimization. You’d have to be insane to not run an NVFP4 quant over a ternary quant on Blackwell if they both fit.