| |
| ▲ | lnenad 3 hours ago | parent | next [-] | | 830t/s is burst aggregate. ~500 is sustained and it's for 8 concurrent users. Meaning for $1.99/hour if you serve 8 users it's 8*$0.54, not just $0.54. You shouldn't rent one out if you're just serving it for yourself, but from a financial standpoint if you sell to users you can take a 100% margin. | | |
| ▲ | minraws 3 hours ago | parent [-] | | Bro 500 Aggregate. so that's 500 * 60 * 60 = 1.8M output which is .5$ at best... Not including pre-fill and stuff. This is not the real margins, even if you are selling to 8 users it's 90 tps per median stream. So assuming that .6-.7$ This is not even remotely worth it. You need to 3x this tps(~1500 tps) to be worth it, and that's what most providers are doing, at 20-30 users at 50-60 tps with better optimized batch processing and kernels you can make some profit. | | |
| |
| ▲ | Almondsetat 3 hours ago | parent | prev | next [-] | | You get privacy for 4 times the cost | |
| ▲ | krisknez 3 hours ago | parent | prev | next [-] | | How is that economically viable? They are selling at a loss? | | |
| ▲ | Aurornis 29 minutes ago | parent | next [-] | | > They are selling at a loss? Definitely not. Inference is not as expensive to operate as many people seem to assume. The frontier labs are probably making a lot of money from selling tokens. It’s covering all of the R&D costs like salaries, collecting training material, and running the large training operations that costs a lot of money. | |
| ▲ | gpugreg 3 hours ago | parent | prev | next [-] | | Agentic workloads are somewhere around 1%/0.5%/98.5% input/output/cached tokens. Cached tokens are pretty much free for inference providers (if they implement sparse and compressed attention properly) and throughput for input tokens is much higher. Lets assume that you've got 2 million input tokens, 1 million output tokens and 98.5 million cached tokens to process. That would cost 2 * $0.14 + 1 * $0.28 + 98.5 * $0.0028 = $0.8358 with DeepSeek API pricing. For comparison, it would take 2M / 8000 + 1M / 800 = 1500 seconds to process this amount of tokens with the linked framework, which is about $0.83 when we assume $2/hr for one MI300X. However, other inference providers have 10 times higher prices for cached tokens, which results in a comfortable margin. And we should not discount that DeepSeek also gets paid in data, which is probably more valuable to them. And I believe that this framework still has some room for optimization for generation with high batch sizes. | | |
| ▲ | throw10920 an hour ago | parent | next [-] | | > should not discount that DeepSeek also gets paid in data, which is probably more valuable to them That's agentic feedback loops for training, right? Any more detail on this, such as how they actually tell whether that data is good or not? That seems like a very hard problem, and like the value of that data is low compared to just building their own, controlled RL gyms. | |
| ▲ | xyzzy_plugh 2 hours ago | parent | prev [-] | | Your math is a bit funny if you're assuming the 1/0.5/98.5 ratios: you doubled input and output tokens but not cached. If you double cached tokens to match your original ratio it works out to around $1.11, and if you 10x the cached token cost it's around $6.08. Based on your $0.83 estimate, the margin isn't great. This is within shooting distance of "at cost" which is probably pretty close to what DeepSeek is operating with, ignoring the value of the data they're collecting of course. > And I believe that this framework still has some room for optimization for generation with high batch sizes. If that optimization can bring this scenario closer to $0.50 then it gets pretty compelling, otherwise I'm not confident. | | |
| ▲ | gpugreg 12 minutes ago | parent [-] | | Oh, I messed up. Half-way through, I thought it would be a good idea to double the numbers so I don't have to deal with half millions, but forgot to also double the 98.5. Unfortunately, I can not edit it anymore. I think the margins of DeepSeek may be a bit better than with this vibe-coded framework here, since they had the liberty of optimizing their models for their own hardware. For DeepSeek V3, they claimed a cost profit margin of 545%: https://github.com/deepseek-ai/open-infra-index/blob/main/20... At the time, open frameworks were not anywhere close to achieving that number. Not sure whether they caught up. The software wizards at DeepSeek are quite skilled. |
|
| |
| ▲ | dietr1ch 3 hours ago | parent | prev | next [-] | | They claim their advantage is knowing how to serve their models efficiently, which is quite possible since they design for it. | |
| ▲ | simlevesque an hour ago | parent | prev | next [-] | | They get all our invaluable data which they'll use to train the next model, to get more data, to train the model after. | |
| ▲ | drob518 an hour ago | parent | prev | next [-] | | I think Deepseek is selling roughly at cost (perhaps a slight premium). They don’t guarantee that they don’t train on the submitted prompts, so I suspect they are mining the data. Mining for what? Well, who knows. Best case, mining to make Deepseek better. That said, I use Deepseek all the time. It has done a whole lot of ‘ls’ commands on my system, though. | |
| ▲ | pama 3 hours ago | parent | prev [-] | | Use nvidia hardware instead and use a larger cluster serving many more users concurrently. Easily 10x–20x higher token rate per GPU with public solutions like dynamo and sglang. |
| |
| ▲ | thrownaway561 3 hours ago | parent | prev [-] | | This is exactly what I came to say. The price of Flash is so cheap that trying to run it locally or with your own hardware is pointless. I was using it about a month ago to program some stuff and ran it for 4 days non-stop and it cost me about $2. | | |
| ▲ | ux266478 2 hours ago | parent | next [-] | | If you don't do any attention steering, custom decoding or meddle with the weights maybe. Services are worthless unless all you do is write positive prompts. As others have mentioned, there's the privacy factor as well. | |
| ▲ | NitpickLawyer 3 hours ago | parent | prev | next [-] | | > trying to run it locally or with your own hardware is pointless. Serving local models has advantages other than price. If you work in restricted industries, or have a strong need to protect your IP, or if you just value privacy more than cost, you now have options. | |
| ▲ | jorvi 3 hours ago | parent | prev [-] | | With the cost of electricity, hardware depreciation and tok/s it rarely makes sense to run locally. |
|
|