| ▲ | margalabargala 30 minutes ago | |
> There is a lot of inefficiency right now keeping prices elevated. That will change very fast and soon paying for inference will likely resemble paying for hosting your app. I don't know about "very fast" or "soon" unless you're speaking in geological terms. SOTA models like Kimi 3 require thousands of GB of RAM/VRAM to run at speeds that are real-time useful. Manufacturing the memory necessary for that quantity to be available at app-hosting prices will take decades. Software efficiency solutions might drop needed memory by an order of magnitude in that time...but a tenth of an enormous amount is still pretty darn big so won't get us there "soon". | ||