Remix.run Logo
villish 20 hours ago

Thats only important if running it locally is critical for privacy reasons or just as a hobby.

Time has a cost in business. If a model needs 30 million tokens to achieve a similar result as another that can do it in 10 million, that 60 tokens per second will take a long time.

hadlock 14 hours ago | parent [-]

Right now qwen 3.6 35b-a3b has a success rate of 92% and qwen 3.8 27b has a success rate of 96%. But the 35b moe does about 1080 tokens/s at concurrency 54, vs 480 tokens/s at concurrency 28. For our specific workflow on blackwell.

Of course enormous batch jobs are different. I was explicit when I said consumer laptop.