Remix.run Logo
Gormo 2 hours ago

We're already at the point where a few months worth of an SMB's token usage from the SaaS LLM providers costs is equivalent to the cost of installing on-prem infrastructure capable of running cutting-edge open-weight models at scale. At my company, we recently installed a server with an array of Gaudi 2 cards on our server rack, and have set up an Open WebUI frontend to expose an LLM connected to all of our internal resources to our staff. The total cost was about $20k, which we'd easily eat up with six months worth of equivalent Claude usage.

We'd originally set this up to be able to locally run larger models, in the 300-400 billion parameter range, but the rate of improvement of open-weight models has been so fast that, coupled with extensive custom skill creation, we're now getting similar results out of Qwen3.8 27B to what we were getting out of Qwen 3.5 297B when we started out with the project, which frees enough memory to allow 20-25 users to have 256K context concurrently. Both the hardware, the software, and the models are improving at an accelerating rate.

Investing in data centers to support SaaS LLM providers today feels a bit like investing in mainframes and minicomputers in the late '70s, with a massive paradigm shift lurking right around the corner.

Actually, it's probably already closer to the early '80s, given that purpose-built local AI workstations are already available at price points lower than the inflation-adjusted initial price of the original IBM PC.