Remix.run Logo
c0m47053 2 hours ago

MoE is great on systems that lack the VRAM to host the full model. On my 16GB VRAM system, I can get 100 tok/s with Q4 Qwen 3.6 35b a3b, and 15 tok/s with 27b.

MTP is a trade-off, as it pushes some more of the model off the GPU.

I have managed to get usable quants of Laguna S2 and even DeepSeek V4 flash on this setup.

There is clearly some intelligence loss compared to similar sized dense models, but I feel like it stomps on the 9-12b models I could run fully on GPU