| ▲ | mdp2021 an hour ago | |
> We're running Kimi 2.8 on a $107k server Equipped with what? Is it CPU based inference, a mix...? | ||
| ▲ | wongarsu 14 minutes ago | parent [-] | |
Not GP, but my educated guess is that they are running a system with between 4 and 6 MI325 or MI355x or similar AMD GPUs. With the 50k tps as the total figure for all parallel requests. Those cards have a lot of memory for their price, allowing you to push to really high batch sizes while still having a large context size for each request | ||