| ▲ | Gareth321 4 days ago | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
> However the cold reality for both is that there is zero moat to a model anymore. The moat right now is a) the hardware, b) the electricity, c) the intelligence, and d) scalability. On hardware, it's very expensive to purchase anything which can provide a fraction of the performance of a subscription. Traditional accounting depreciation would imply that purchasing local hardware is a terrible financial decision. On electricity, this is a surprising cost center depending on location. A system with just one 5090 can easily pull 1kW, and to achieve usable performance for a workplace is going to require dozens of machines. This can represent an extra $10-20k in electricity in cheap places. In California or Europe this could be $30-60k per year. As for intelligence, the frontier models from OpenAI and Anthropic are still superior, and they have at least a 3-6 month head start. Distilled models are closing the gap on some metrics, but they still can't compete. That's why they cost so much less. The last major moat is the ability for subscriptions to scale with need. This means easily adding and removing licenses. This is far easier than purchasing extremely expensive hardware (and managing it), and selling it if/when internal demand changes. It's the same reason companies use contractors. The ramp up/down costs are very high. The only real moat that local LLMs have right now is privacy. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | jurgenburgen 3 days ago | parent | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I think you went from one extreme to another. OpenAI and Anthropic rent their compute from AWS & friends. When we say large enterprises are moving to open weight models it means they are cutting out the middleman and renting the compute directly from AWS instead of giving OpenAI and Anthropic a margin. > they have at least a 3-6 month head start This is a moat of nothing. Our company still hasn’t gotten access to Fable so switching to open weight models would mean getting access to similar quality models. In some orgs they are still on 2025 models. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | sedansesame 4 days ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Privacy is non-negotiable for corporate. Even without considering costs or country of origin, we've seen from OpenAI that claims of AI safety are worth less than the (virtual) paper they're printed on. All it takes is one incident, and all your company's internal data will start showing up in public users' chats. You can rely on a contract to prevent this, or you can guarantee it by using a locally hosted model you fully control. When combined with the cost savings and good enough performance mentioned in the article, this can become a huge selling point. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | dingaling 3 days ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
"Traditional accounting depreciation would imply that purchasing local hardware is a terrible financial decision." Depreciation is designed to _encourage_ purchasing of useful local tools, by incrementally matching fractions of the cost of the tool to the revenue it generates over its useful life. The fact that a graphics card might have a book value of $0 after five years of depreciation is a feature, not a bug. Since the invention of corporation tax it has also had the benefit of offsetting tax over the same period, instead of just one big offset in the first year. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | archagon 4 days ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Actually, the moat is regulatory. Expect these companies to behave themselves in progressively more grotesque and sycophantic ways to get the federal government to make open/foreign models (and their output) illegal. After all, their very survival depends on it. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | cmiles8 4 days ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
You’re missing the point that 98+% of the use cases for AI don’t require the latest greatest model and are far better positioned to use the fast-follow distilled cheap models. OpenAI and Anthropic are fighting to win a race (build the biggest baddest model) that has no prize. The prize is mass adoption at scale at the best price, which is why companies are rapidly shifting to open model. They don’t need to pay 10x for a model that’s provides no practical additional benefit. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | tancop 3 days ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
You don't need to self host to get the benefits of an open model. There are many hosted providers cheaper than OpenAI or Anthropic who can give you a SLA, ZDR, BAA and all the other three letter acronyms your compliance department needs. The important part is if they break the contract or raise their prices you can always move to a different provider. You get lower cost and lower risk at the same time which is extremely rare in business. That's just not possible for closed models where your only options are the official branded API or Azure/Bedrock. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | wolvoleo 3 days ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Hmm yes but the equipment cost moat is artificial. This scarcity was created by the big AIs by buying up all the future production capacity. That works for a while but it won't last forever. It's the same with the subscriptions. Local models can't compete because they're simply giving too much value for money. They're effectively subsidised by Big AI. Again something that won't last. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | subarctic 4 days ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Re this point > As for intelligence, the frontier models from OpenAI and Anthropic are still superior, and they have at least a 3-6 month head start. Distilled models are closing the gap on some metrics, but they still can't compete. That's why they cost so much less. I would argue that the reason they cost so little is because anyone can run open models and offer them as a service, so there's actual competition and the price is closer to cost. i.e. if the open models were just as intelligent as frontier models but cost the same to run as they do right now, the price wouldn't be higher (unless demand went up so high that marginal cost to provide more of the service went up, due to scarcity of hardware and or electricicy). On the other hand, if what you're saying is the frontier labs have some pricing power due to their models being better, and that is the reason they are able to charge more than the companies providing open models as a service, then I would agree. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | jjav 3 days ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
> As for intelligence, the frontier models from OpenAI and Anthropic are still superior I'll grant they are superior at least right now. But also, they are too expensive. We ($work) are finding that it is best to build engineering discipline around AI usage (who would've thought!) and use the cheaper models like Cursor Composer. Using Opus we can blow through an entire month budget in an afternoon, so while more powerful, it is no longer practical except for rare very complex tasks. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | alxfoster 3 days ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Your annualized energy cost estimates are off by an order of magnitude. 1kwH @ $0.1 (Texas) is $2.40/day if 100% utilized 24x7, California is roughly twice that per my understanding. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | xoa 4 days ago | parent | prev [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
>On hardware, it's very expensive to purchase anything which can provide a fraction of the performance of a subscription. One of the basic questions/concerns here though is that it's not like the AI places are getting the GPUs for 10x less. It's true they have some economies of scale, but they also have some waste, and frankly in this particular case it's not clear they get that much gain over what a lot of businesses could achieve. The biggest traditional gain for central providers is that a lot of typical computing usage is burst-y, and in turn local kit might be underutilized. But with LLMs heavy users tend to use them all the time assuming their tokens allow it (and in the case of local hardware there's nothing stopping you, quite the contrary), they can use it directly interactively or leave them to go overnight on something too. So it's reasonable to suspect that the reason subscriptions are only a fraction of the cost is that we're in a bubble seeing these companies losing money in an attempt to gain some sort of durable advantage. Just as every previous time, there is the chance that the music stops at some point, and they need to crank up pricing or pull other schemes to actually make money. Of course, it can be a good deal in the mean time, you basically get to suck down investor money for nothing, but it's also not unreasonable to at least be consider fallbacks. Even beyond questions of control and risk etc. I know at least a few places that are now genuinely considering questions like "what happens if a datacenter we depend on gets droned" that would have never had an iota of thought devoted to them even 5 years ago. >On electricity, this is a surprising cost center depending on location. A system with just one 5090 can easily pull 1kW, and to achieve usable performance for a workplace is going to require dozens of machines. This can represent an extra $10-20k in electricity in cheap places. I don't think that's "surprising" at all, everyone knows about power use. And this seems like it gets heavily into what you're defining as "usable" and is also more useful to define in terms of cost-per-employee vs total. Obviously a bigger business will have a higher line number total even if the cost per employee is identical, but simultaneously can be expected to be making more revenue to pay for it. If we're defining an average of a dedicated 5090 pulling 1 kW for every single employee (presumably some people wouldn't use it all the time, but others would then pull the compute for other work), running 24/7 (to cover people running stuff when they're away), then that'd be 8760 kWh per year. At my not particularly cheap New England location that'd be about $1900 per employee per year at the generalized residential rate (~$0.22/kWh), or $156 per month. That doesn't seem radical if it really does boost productivity. However, there is a lot of room to go lower. I'd expect a business to run backup anyway, and these days there are a lot of incentives to do that at least partially with batteries. That also opens up rate shifting as another way to pay back the cost. If we change to time of day pricing, that's 8 hours of peak pricing with the rest off-peak. 8 kWh of battery can now be had for a few thousand. And the off-peak rate is only ~$0.14/kWh, cutting the cost per year by about $700 to $1200 per employee per year. Solar power is also usually far more valuable to use yourself then sell back to the grid, and also continues to plummet in price. None of this is to say that it makes sense for every place at all, but it's close enough to the the line that the math is at least worth exploring, or could at least lower the cost enough to be worth it given other things. It really comes down to how much extra value the company (or individual) expects to come out of it per month. >In California or Europe this could be $30-60k per year. Dunno about Europe, but at the kinda prices I see for California I'm really surprised more places aren't trying to move a lot of usage to battery+renewable. >The only real moat that local LLMs have right now is privacy. I don't think resiliency and control are things that can be taken for granted anymore, particularly on the global scale. War and terrorism is getting worse again. International relations are getting nastier, and governments have the power to just order places cut off. If LLMs aren't particularly valuable to a business, then why an expensive subscription? But if they are particularly valuable, then insurance is something leadership should be contemplating. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||