| ▲ | kees99 4 hours ago | |||||||||||||||||||||||||||||||||||||
Qwen models are slower in tokens/s, compared to similarly sized gemma4 and others, and they use more tokens per task, in part thanks to that xhigh default. On the other hand, there are some of us who are stuck with hardware that has plenty of compute, but limited (V)RAM. The new 27B is just perfect for that. | ||||||||||||||||||||||||||||||||||||||
| ▲ | petu 4 hours ago | parent [-] | |||||||||||||||||||||||||||||||||||||
> Qwen models are slower in tokens/s, compared to similarly sized gemma4 and others No? Gemma 31B and Qwen 27B are about the same speed. Gemma 26B-A4B and Qwen 35B-A3B are about the same speed. | ||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||