| ▲ | toomuchtodo a day ago | |
My primary role is cybersecurity in a regulated entity in a regulated industry, I am highly confident it is straightforward to do so based on work accomplished in only a couple of weeks. Stand up a router, stand up a Kubernetes cluster if you don't have one, stand up the necessary VMs and compute for serving inference. Two pizza team, in my experience. Customers can switch (although we can argue the speed and pain of doing so), and the speed at which they do will be a function of cost efficiency and demonstrable value (imho). A recent example of this is Broadcom and VMware [1], for example. When motivated, it can be done. If there is no objective, measured value being delivered, the spend will be cut. If the value delivered is measured, it will be enabled at a lower cost through cost optimization measures (ie self hosting) [2]. This is all to say: there is no moat, the revenue of inference providers is volatile and not assured in any measure. Caveat emptor. [1] https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu... [2] Microsoft considers replacing ChatGPT and Claude with Kimi K3 to save $600M - https://news.ycombinator.com/item?id=49022984 - July 2026 | ||