Remix.run Logo
moshun an hour ago

Considering the rate of model development and rail hopping, seems like baking models into silicon is speed-running obsolescence.

breuleux an hour ago | parent | next [-]

If you’re only running models for frontier capabilities, yeah. For tasks where current models are smart enough, running them 100x faster is the most impactful improvement you can make. Consider all the things you could use a model for, but don’t, because the latency is just a bit too high.

zxspectrum1982 27 minutes ago | parent | prev | next [-]

I'd gladly pay for a Claude Opus 4.6 Thinking High in silicon and use it for 1-2 years. It's good enough for many coding tasks.

Gigachad 2 minutes ago | parent [-]

It costs something like $300,000 for the hardware to run a model of that size. You'd pay that for a single model for 1-2 years? Not even the AI companies can justify that kind of spend which is why they keep extending the expected lifespan on their hardware in the accounting.

zxspectrum1982 a minute ago | parent [-]

I'm expecting the Taalas MSIC version to cost a fraction of that. Then probably have some kind of cheap subscription to Anthropic for updates (yes, Taalas chips can receive a certain kind of updates: they have a small SRAM).

mdp2021 an hour ago | parent | prev | next [-]

Compute the cost of producing n of them devices, imagine a fair price based on that, and see if that local, blazing fast card* can be an asset that could be replaced periodically.

*(It's local: private files managing firm oriented. It's blazing fast: it can be placed into recursive, intensive local workflows.)

try-working 27 minutes ago | parent | prev | next [-]

obsolescence is the whole point. apple gets to sell a new phone very 6-12 months because of it.

i have written about this:

"For device makers

Packaging models with laptops and smartphones will let application access near free, low latency inference and potentially offer users a better experience with the option of preserving data on-device. This is viable under the condition that tasks that do require larger expert models that run in the cloud can be routed to external models. A side-effect of local models and what will let Apple cut upgrade cycles from ~4 years (?) down to 12-18 months is specialized hardware to run them. For almost a decade, smartphones have been trying to compete on better cameras. This coming decade will see them selling better GPUs, NPUs, ASICs and whatever other things they'll be calling the inference chips, to drive re-purchase. Every six months will see a better model on new hardware, which will enable better performance in certain applications."

https://try.works/role-model-the-case-for-a-model-routing-pr...

nomel 6 minutes ago | parent [-]

No, the point is inference speed and power.

topspin an hour ago | parent | prev | next [-]

"seems like baking models into silicon is speed-running obsolescence"

Now maybe. When models are flying passenger aircraft, other prerogatives will assert themselves. When a 50TB ROM means you can impulse purchase a ChatGPT 6.3 xhigh that runs on batteries, yet more use cases will be apparent.

mdp2021 38 minutes ago | parent [-]

Well, 50TB ROM Taalas HC1 style would be apparently a 400000b transistor system through a chip sized 2.5 meters on the side... :)

thfuran 5 minutes ago | parent [-]

Phones were getting too thin anyways.

ray_v an hour ago | parent | prev | next [-]

I could see this making sense when model development start to settle down ... it's going to settle down, right? ...

alightsoul an hour ago | parent | prev | next [-]

Which is exactly what companies and shareholders want to increase sales.

amelius an hour ago | parent | prev | next [-]

Not sure. You can fix the transistors but leave the connections between them open for flexibility, so you only need to change the manufacturing process for the upper masks for every new model.

tsujamin an hour ago | parent | next [-]

Surely that added flexibility negatively impacts the density/parameter count of the model you could etch?

sroussey an hour ago | parent | prev [-]

Or do a hybrid

flyinglizard an hour ago | parent | prev [-]

Look at it the other way: compared to the cost of training a model, the cost of making a custom ASIC is trivial.