| ▲ | cmrdporcupine 3 hours ago | |
Those two offer MoE variants, this doesn't seem to. Dense model makes it dog slow on anything without HBM. Max 15tok/sec on decode on DDR5 systems like a Spark or a Strix Halo -- and that's at 4 bit quant. | ||
| ▲ | EddieRingle 2 hours ago | parent | next [-] | |
Dense models run at a very usable speed (Qwen 3.6 was running at ~50t/s last I looked) on my dual 7900 XTX desktop. (And before anyone brings it up, I did not buy them for this purpose, so the up-front cost is irrelevant in my case.) | ||
| ▲ | Havoc an hour ago | parent | prev | next [-] | |
The benchmark comparison is against the dense variants not MoE | ||
| ▲ | petu 3 hours ago | parent | prev [-] | |
3090/4090 probably would do 40 t/s, for 5090 75 t/s is shown in the blog. | ||