Remix.run Logo
sschueller 3 hours ago

Model-on-Chip is coming. GPU are for general computing but have a huge bottle neck for doing model inference.

Even not being able to significantly update a model that is burned on a chip the performance gains are immense. You also don't need the latest chip fabs to make them drastically reducing the cost.

thisoneisreal 2 hours ago | parent | next [-]

One thought I had is that you could use FPGAs to get hardware performance but maintain the ability to dynamically update. I don't know enough about hardware to consider trying such a thing but I'm curious if that could be made practical and economical somehow.

baby_souffle 2 hours ago | parent | prev | next [-]

I don't think asics specific to a specific model or even model family are likely to be commodity hardware anytime soon.

It's extremely expensive to build that and you'll be at least two major model generations behind before you even get your first wafers back. By the time you got your production run ready to go and packaged for market nobody's going to care.

Once we end up going something like 24 months between major advances and capabilities for these models then I can start to see asics for a model being possible.

cousinbryce an hour ago | parent | prev | next [-]

This will probably work well with SotA planning and local chip implementation. I see them being like cars. Cost a few 10k on credit, buy a new one when the old one goes bad or marketing convinces you to upgrade.

nicce 3 hours ago | parent | prev | next [-]

Many years until consumers can buy them at reasonable price. Nvdia and AMD are making GPUs bad in purpose for consumers so that nobody can build a datacenter from them. It will take a long time.

Danox an hour ago | parent [-]

Hardware and software getting better every day the barbarians are at the gate…

ben_w an hour ago | parent | prev [-]

> Even not being able to significantly update a model that is burned on a chip the performance gains are immense. You also don't need the latest chip fabs to make them drastically reducing the cost.

Yeeeees but the models are in some sense doubling in performance every 4 months, so I expect this to happen in serious quantities approximately when the economic bubble bursts and investors are no longer willing to pay for training.

(Based on widespread news reporting of the existing impact on US electricity markets, I expect this around the end of this year; but with regards to news reporting I am aware of the Gell-Mann amnesia effect, so if this is as much BS as the water issue turned out to be…)