Remix.run Logo
FuriouslyAdrift 3 hours ago

We're running Kimi 2.8 on a $107k server and getting around 50k tokens per second or better on most things.

We've already saved money compared to last years token cost on Claude/Gemini

mdp2021 an hour ago | parent | next [-]

> We're running Kimi 2.8 on a $107k server

Equipped with what? Is it CPU based inference, a mix...?

wongarsu 16 minutes ago | parent [-]

Not GP, but my educated guess is that they are running a system with between 4 and 6 MI325 or MI355x or similar AMD GPUs. With the 50k tps as the total figure for all parallel requests. Those cards have a lot of memory for their price, allowing you to push to really high batch sizes while still having a large context size for each request

ajju an hour ago | parent | prev | next [-]

What is your config?

freediddy an hour ago | parent | prev | next [-]

how many requests per second can the server take?

logicallee an hour ago | parent | prev [-]

Are you developing software? Is most of it used on a coding agent? (Like Claude Code or ChatGPT Codex?) If so, what coding agent do you use? If you're not developing software what do you use it for (roughly)?