| ▲ | pllbnk 7 hours ago | |
Just a couple days ago I learned about ninfer (https://github.com/Neroued/ninfer) and on RTX 5090 I can now get ~200 tok/s and over 400 tok/s on concurrent requests which is plenty fast for a local model of this strength. | ||
| ▲ | jakswa 3 hours ago | parent | next [-] | |
dang only for certain nvidia GPUs, had my hopes up | ||
| ▲ | lowbloodsugar 3 hours ago | parent | prev | next [-] | |
Ok, I need to try that. I'm getting 45tok/s with vLLM on my 6000. >600tok/s concurrent, but 45tok/s single request. | ||
| ▲ | beastman82 7 hours ago | parent | prev [-] | |
can't second ninfer enough. amazing tech | ||