| ▲ | c0m47053 2 hours ago | |
MoE is great on systems that lack the VRAM to host the full model. On my 16GB VRAM system, I can get 100 tok/s with Q4 Qwen 3.6 35b a3b, and 15 tok/s with 27b. MTP is a trade-off, as it pushes some more of the model off the GPU. I have managed to get usable quants of Laguna S2 and even DeepSeek V4 flash on this setup. There is clearly some intelligence loss compared to similar sized dense models, but I feel like it stomps on the 9-12b models I could run fully on GPU | ||