| ▲ | ernsheong a day ago | |
Yes it's between this and Gemma 4 31B which is much slower, but looks like it won't ever get an upgrade. I have to conclude that the MoE variants are unreliable, and MTP sometimes just can't get tricky formatting right. | ||
| ▲ | dofm a day ago | parent | next [-] | |
The whole series had an upgrade a couple of days ago actually — they have addressed embedded tool calling (and hopefully the MTP formatting stuff though I gave up running the Gemma MTP because it's often slower than not-MTP) Not tried it yet but I've seen tests that suggest they've properly fixed the tool calling issues. | ||
| ▲ | SwellJoe a day ago | parent | prev | next [-] | |
I find the 4-bit QAT with MTP to be entirely usable speed on both my boxes (Strix Halo and a desktop with two V620 GPUs, which are slightly faster than the Strix Halo). | ||
| ▲ | cmrdporcupine a day ago | parent | prev [-] | |
For whatever reason prefill (on my DGX Spark) is faster with the Gemma models than Qwen 3.6 models of similar size. On vLLM anyways. Likely just deeply tuned code contributed to vLLM by Google? vLLM gives me ~7000+ tok/sec with Gemma 4's MoE model. Vs ~6000 tok/sec for Qwen 3.6 MoE. | ||