Single 3090 under llama.cpp:
| model | size | test | t/s |
| ------------------- | ------- | ------ | ---- |
| gemma4 31B Q4_0 | 16.1 GB | pp2048 | 1248 |
| gemma4 31B Q4_0 | 16.1 GB | tg512 | 40 |
| qwen35 27B Q4_K | 15.9 GB | pp2048 | 1248 |
| qwen35 27B Q4_K | 15.9 GB | tg512 | 39 |
| gemma4 26B.A4B Q4_0 | 13.3 GB | pp2048 | 4304 |
| gemma4 26B.A4B Q4_0 | 13.3 GB | tg512 | 160 |
| qwen35 35B.A3B Q3_K | 15.7 GB | pp2048 | 3329 |
| qwen35 35B.A3B Q3_K | 15.7 GB | tg512 | 144 |
> with their respective speculative decoding methodsYou're benchmarking drafter acceptance rate, then. Which is real life values, yes, but attributing worse drafter performance to the other 95% of the model being inherently slower.