| ▲ | seanmcdirmid 2 days ago | ||||||||||||||||
I get up to 90 tokens / second with Jundot/Qwen3.6-35B-A3B-oQ6-mtp, on a M3 Max with 64GB. MoE so it is not a dense model, but that means it runs faster (also, mtp helps). It is a 30GB model, but you should be able to load it, otherwise try the 4-bit quant instead of the 6-bit quant, don't bother quanting your KV Cache (don't enable turboquant in oMLX), since that will slow you down. I'm not sure what that means on a M5 max, definitely faster, I don't know if it really plays into the strengths of the new chip design though. | |||||||||||||||||
| ▲ | coder-pm a day ago | parent [-] | ||||||||||||||||
Thanks! I have to try that! Might be tight! Can it run in Claude Code? Are you loading it with Ollama? | |||||||||||||||||
| |||||||||||||||||