| ▲ | Foobar8568 an hour ago | |
Well, according to claude and Jevbench, Qwen 3.6 35b with ninfer on a RTX 5090@480W is like 3-5 time slower but 10%-15% better performance on the public set, I could see prefill > 15k for 700-800decode. Latency against what and which hardware? I don't really get jev... | ||
| ▲ | alex7o an hour ago | parent [-] | |
Look I can convince my boss to pay for jev, but I won't convince him to run our prod stuff on a rented vast.ai 5090. And the pricing wouldn't be worth it. If you have ideas I would be glad to hear them | ||