| ▲ | Ilaurens 3 hours ago | |||||||||||||||||||||||||
These token numbers look off for a 5090. For comparison, an rtx 5000 Blackwell SFF with just 470gb/s bandwidth gets me 30 tok/s on gemma4-31B-QAT. Almost 60 tok/s with MTP enabled. A 5090 should get you much more than that! | ||||||||||||||||||||||||||
| ▲ | jszymborski 3 hours ago | parent | next [-] | |||||||||||||||||||||||||
Thanks for flagging, this is on my local rig and it's driving my display too. I'm curious now, will take a closer look. These are the tok/s as reported by LMStudio. EDIT: Updating llama.ccp gets me 58 tok/s on Gemma 31b | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||
| ▲ | mft_ 3 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||
Agree - I got 9.7 tok/s on an M1 Max with unsloth's gemma-4-31B-it-qat-UD-Q4_K_XL. | ||||||||||||||||||||||||||
| ▲ | 3 hours ago | parent | prev [-] | |||||||||||||||||||||||||
| [deleted] | ||||||||||||||||||||||||||