| ▲ | jszymborski 3 hours ago | ||||||||||||||||
Thanks for flagging, this is on my local rig and it's driving my display too. I'm curious now, will take a closer look. These are the tok/s as reported by LMStudio. EDIT: Updating llama.ccp gets me 58 tok/s on Gemma 31b | |||||||||||||||||
| ▲ | fhars 2 hours ago | parent [-] | ||||||||||||||||
The Qwen-35B-A3B numbers are even weirder, did you drop a digit? I get half of that speed on a AMD Radeon RX 5500 XT (RADV NAVI14) (8192 MiB) (unsloth/Qwen3.6-35B-A3B-MTP-GGUF:Q6_K with q8_0 context). | |||||||||||||||||
| |||||||||||||||||