Remix.run Logo
fithisux a day ago

I would be more interested the cost of each model in hardware resources to achieve 1 token per second.

Cost is a metric of interest these days.

adrian_b 21 hours ago | parent [-]

The cost is an extremely nonlinear function of the speed expressed in tokens per second.

Throughputs of around 1 token per second can be achieved even for the biggest currently existing open-weights models with a mini-PC with 2 internal M.2 SSDs, which by the beginning of this year would have cost around $1200, though now it probably costs more than this, because the prices of DRAM & SSDs have continued to escalate.

It would be possible to increase the speed about 4 times for a double cost, while still running the model from SSDs in a desktop with extra SSDs in an add-on card, in a 16-lane PCIe 5.0 slot.

Then to increase more the speed you would have to fit all the weights in DRAM, and at the current prices 1 TB or 2 TB of DRAM in a server computer would increase the cost more than one order of magnitude.

For more speed, the weights must fit inside the GPU memory, which would bring again a cost increase of more than 1 order of magnitude.

So the cost per token per second is meaningless due to the great non-linearity of the cost, as a function of both model size in GByte and of throughput in tokens per second.