| ▲ | wongarsu 19 hours ago | |
While running the model at 9000 tokens/s is the more flashy demo, I imagine running 1000 concurrent requests at 150 tokens/s each is the much more achievable goal | ||
| ▲ | LarsDu88 17 hours ago | parent | next [-] | |
At 9000 tokens/s you could interleave a lot of requests so long as pre-fill is also fast. It really depends on how much you need to keep sessions open to take advantage of KV caching | ||
| ▲ | wyre 18 hours ago | parent | prev | next [-] | |
It depends. If I was running a model locally I would much prefer 9000 t/s. If I was running an inference company, obv 1000 concurrent requests at 150 t/s is preferable. | ||
| ▲ | OliverGuy 18 hours ago | parent | prev [-] | |
[dead] | ||