| ▲ | petu 4 hours ago | |||||||||||||
> Qwen models are slower in tokens/s, compared to similarly sized gemma4 and others No? Gemma 31B and Qwen 27B are about the same speed. Gemma 26B-A4B and Qwen 35B-A3B are about the same speed. | ||||||||||||||
| ▲ | trouve_search 3 hours ago | parent | next [-] | |||||||||||||
What configuration are you using? On both vllm and llama-cpp, I get significantly higher speeds from gemma4 than qwen3.6 (with their respective speculative decoding methods). Output TPS in vllm for instance: - Gemma4 26B-A4B: 200-300TPS - Qwen3.6 35B-A3B: 120-180TPS - Gemma4 31B: 80-120TPS - Qwen3.6 27B: 60-80TPS This is for a first request on a dual 5090 setup, with their respective speculative decoding methods. | ||||||||||||||
| ||||||||||||||
| ▲ | stymaar 3 hours ago | parent | prev [-] | |||||||||||||
There's no Qwen3.8-35B-A3B though. | ||||||||||||||
| ||||||||||||||