| ▲ | trouve_search 3 hours ago | |
What configuration are you using? On both vllm and llama-cpp, I get significantly higher speeds from gemma4 than qwen3.6 (with their respective speculative decoding methods). Output TPS in vllm for instance: - Gemma4 26B-A4B: 200-300TPS - Qwen3.6 35B-A3B: 120-180TPS - Gemma4 31B: 80-120TPS - Qwen3.6 27B: 60-80TPS This is for a first request on a dual 5090 setup, with their respective speculative decoding methods. | ||
| ▲ | xfalcox 5 minutes ago | parent | next [-] | |
Have you tried running it on a single 5090? Dual 5090 require https://github.com/aikitoria/open-gpu-kernel-modules for higher perf. Are you using TP? | ||
| ▲ | petu 2 hours ago | parent | prev [-] | |
Single 3090 under llama.cpp:
> with their respective speculative decoding methodsYou're benchmarking drafter acceptance rate, then. Which is real life values, yes, but attributing worse drafter performance to the other 95% of the model being inherently slower. | ||