| ▲ | villish 20 hours ago | |
Thats only important if running it locally is critical for privacy reasons or just as a hobby. Time has a cost in business. If a model needs 30 million tokens to achieve a similar result as another that can do it in 10 million, that 60 tokens per second will take a long time. | ||
| ▲ | hadlock 14 hours ago | parent [-] | |
Right now qwen 3.6 35b-a3b has a success rate of 92% and qwen 3.8 27b has a success rate of 96%. But the 35b moe does about 1080 tokens/s at concurrency 54, vs 480 tokens/s at concurrency 28. For our specific workflow on blackwell. Of course enormous batch jobs are different. I was explicit when I said consumer laptop. | ||