| ▲ | NooneAtAll3 16 hours ago | |
I don't understand the premise in the beginning how is running servers supposed to be 0 cost, while running ai inferrence isn't? | ||
| ▲ | cheema33 14 hours ago | parent | next [-] | |
> how is running servers supposed to be 0 cost, while running ai inferrence isn't? For a SaaS business, running servers isn't free. But compared to the cost of running GPUs for inference that you are selling, it almost is. The company I work for is a SaaS company. We have a single production server. A couple of QA servers. All hosted on Hetzner. Monthly cost for servers is less than $400. This generates a few million dollars a year in revenue. If we were in the business of selling inference, our cost of providing the service, for the same amount of revenue would significantly higher. Even large businesses like Microsoft, Meta, Google have operated with similar margins. Cost of running servers, compared to revenue was very low. But inference changed that, in a dramatic way. | ||
| ▲ | throwawayffffas 16 hours ago | parent | prev [-] | |
A typical server that costs 10k to 30k to own and operate can serve between hundreds and thousands of requests per second of a traditional web application like facebook for 2-4 kW of power, the marginal cost of each request is effectively zero. A single response from kimi k3 requires hardware that cost between 500k and 1m dollars up front and draw over 20kW. Each request costs at least 5% to 10% of the charged cost. | ||