| ▲ | mirekrusin 3 hours ago | ||||||||||||||||
Personally I find speculative decoding much better strategy than MoE – performance wise it's there at 90-100 t/s on 2x4090, great intelligence – really great fit. | |||||||||||||||||
| ▲ | d4rkp4ttern 3 hours ago | parent | next [-] | ||||||||||||||||
A lot of people, including me, don’t want to bother with GPUs, they’d rather run it on their M1-M5 MacBook. For example the 35B-A3B is very usable even on a M1 64GB MacBook. | |||||||||||||||||
| |||||||||||||||||
| ▲ | c0m47053 2 hours ago | parent | prev [-] | ||||||||||||||||
MoE is great on systems that lack the VRAM to host the full model. On my 16GB VRAM system, I can get 100 tok/s with Q4 Qwen 3.6 35b a3b, and 15 tok/s with 27b. MTP is a trade-off, as it pushes some more of the model off the GPU. I have managed to get usable quants of Laguna S2 and even DeepSeek V4 flash on this setup. There is clearly some intelligence loss compared to similar sized dense models, but I feel like it stomps on the 9-12b models I could run fully on GPU | |||||||||||||||||