| ▲ | LaurensBER 2 hours ago | ||||||||||||||||||||||||||||
Dax (from Opencode) has tweeted that they can replicate or beat the price with rented GPUs. Deepseeks secret sauce is the incredibly cheap caching (magnitude cheaper than other providers). vLLM has recently released a similar approach. It's not as effective as what DeepSeek does but still an interesting development. I have no doubt that in due time other providers will match or perhaps even beat the current DeepSeek prices. | |||||||||||||||||||||||||||||
| ▲ | minraws an hour ago | parent | next [-] | ||||||||||||||||||||||||||||
As someone who recently tried it on some blackwell cards, it's possible to match the prices especially the input can be even cheaper and output can match the costs so you can easily build a net 20-30% margin business even at current GPU prices. The entire issue is caching, I tried to write some custom to dump to disk kv-caching using some ideas from their papers and my experience with snapshots and vm checkpoint systems, I must say they must have really squeezed that lemon it's hard. Atleast me with Sol couldn't figure it out over a couple days, a few hours each day, which isn't much but I did feel a bit stuck with existing solutions and felt like I might have to write something from scratch. But if you are willing to put in the effort into the infra I do think it's doable. But it will be really hard to pull it off. My congrats to anyone who manages to pull it off, they might be able to kill off most AI labs. Assuming they can find the compute, Deepseek really has killed all models for me other than Sol/Fable/Opus/K3 tier stuff. | |||||||||||||||||||||||||||||
| ▲ | twotwotwo an hour ago | parent | prev | next [-] | ||||||||||||||||||||||||||||
One read is 1) they're getting a lot of traffic for Flash, 2) they've said they're updating Pro soon and expect that to lead to a traffic spike for Pro, but 3) that would leave them overloaded, so 4) they're going to raise prices to avoid it. It's interesting that most open models adding 1M context did it in a way that reduces KV cache size (though DeepSeek was the most aggressive, using compressed attention on all layers), but only a couple providers turned it into a discount on cache reads. | |||||||||||||||||||||||||||||
| ▲ | _aavaa_ 19 minutes ago | parent | prev | next [-] | ||||||||||||||||||||||||||||
I'll believe it when I see it. Their prices are still much higher than deepseek, especially the caching. | |||||||||||||||||||||||||||||
| ▲ | NorwegianDude an hour ago | parent | prev | next [-] | ||||||||||||||||||||||||||||
Eh, what are you guys even talking about? Deepseek is not cheapest provider as is, and it's MIT. So deepseek making it more expensive to use is just nonsense, they can only change their own pricing. It's the beauty of MIT license and open weights. If anything, these models are some of the safest in the world to use if you worry about a rug pull. | |||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||
| ▲ | retinaros an hour ago | parent | prev | next [-] | ||||||||||||||||||||||||||||
any link to this caching tech? | |||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||
| ▲ | onlyrealcuzzo an hour ago | parent | prev [-] | ||||||||||||||||||||||||||||
> Deepseeks secret sauce is the incredibly cheap caching (magnitude cheaper than other providers). Can anyone working at one of the main US labs (Google, OpenAI, Anthropic) comment on WTF they haven't even tried MLA - despite the obvious massive advantages? I know enough to know they aren't completely incompetent. So there must be a quite good reason. But it remains a mystery to me. DeepSeek's MLA is like almost 2 years old at this time. They've got thousands of people working on this stuff. They clearly have the ability to at least try it... | |||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||