Remix.run Logo
puzzlingcaptcha 3 hours ago

What sort of pp/tg speed do you get on a Strix Halo?

cpburns2009 2 hours ago | parent | next [-]

This is the best I got, all with Unsloth's quantizations.

Laguna-S-2.1:UD-Q4_K_XL (no MTP) pp=186.4 t/s tg=27.8 t/s

Qwen3.6-35B:UD-Q4_K_XL (with MTP) pp=404.4 t/s tg=83.2 t/s

Qwen3.6-27B:UD-Q4_K_XL (recorded pre-MTP) pp=343 t/s tg=12.1 t/s

Laguna actually performed better than I remembered. I thought it was slower.

an hour ago | parent [-]
[deleted]
ascii0eks84 an hour ago | parent | prev [-]

What are pp/tg? I get 30t/s on 27B qwen.

throwawayffffas an hour ago | parent [-]

pp is prompt processing how fast it processes the prompt. Tg is token generation how fast, it generates tokens.