| ▲ | LarsDu88 2 hours ago |
| I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference. |
|
| ▲ | anthonypasq an hour ago | parent | next [-] |
| Personally I think Apple should have acquired them. if you could burn a gemma4 class model into an iphone and actually get extremely low latency and low battery usage it would feel like the future IMO. even if it means you wont get frontier intelligence, there might actually be incentive to buy a new mobile device every year again. |
| |
| ▲ | Melatonic 38 minutes ago | parent | next [-] | | The Taalas chips are not physically small. And part of their secret (if you look at the design) is just locating a bunch of memory soldered on the edges ( I belive higher amounts of SRAM ? ) | |
| ▲ | adgjlsfhk1 an hour ago | parent | prev | next [-] | | I don't think this works out from a cost/silicon perspective. Small models already run pretty well in software (since the weights fit in cache) and big models require silicon area proportional to the size of weights. On a mobile device putting a chip like this is competing directly in BOM and power against a whole lot more l3 cache, and the l3 cache makes everything faster | | |
| ▲ | bastawhiz 32 minutes ago | parent | next [-] | | The weights might fit in cache, if you're using a small model. If you wanted to have a 20B+ parameter model, that's just going in RAM. You could put more RAM in the device and pay the perf cost or have a dedicated chip. Most devices already have a dedicated chip, this just changes which silicon you're spending the money on. | |
| ▲ | teaearlgraycold 35 minutes ago | parent | prev [-] | | My question is what changes about LLM use cases when you’re getting 1000 tok/s? Models in silicon might dramatically change how we think about them. | | |
| ▲ | RussianCow 28 minutes ago | parent [-] | | That likely isn't as relevant for on-device iPhone usage as it is for Real Work™. I won't notice the difference between 50tps and 1000tps when asking Siri a question. |
|
| |
| ▲ | bsaul an hour ago | parent | prev | next [-] | | That's actually a really good point... There's currently zero incentive to buying more hardware, and that's one very good reason do have a new one. | |
| ▲ | superb_dev an hour ago | parent | prev | next [-] | | From what I remember, these chips are not mobile size yet | | |
| ▲ | bradfa an hour ago | parent [-] | | A small model would be. I think that’s more the point. It’s definitely not SOTA but it’s fast and energy efficient and local. | | |
| ▲ | mdp2021 44 minutes ago | parent | next [-] | | > A small model would be [mobile size] A ~30mm side for the HC1 tech for an 8b model (still unclear the planned HC2)? | | | |
| ▲ | wmf 41 minutes ago | parent | prev [-] | | Nope, a small model would be larger than the whole iPhone SoC. |
|
| |
| ▲ | whatsThisBtn4 37 minutes ago | parent | prev [-] | | Apple is somewhere between fashion company and second rate tech company. They could have 9 year old AI and still post profits. Not sure if it's my pixel or android, but I made a randos jaw drop with what the crappy AI on android can do. When are we getting android OpenClaw? |
|
|
| ▲ | moshun an hour ago | parent | prev | next [-] |
| Considering the rate of model development and rail hopping, seems like baking models into silicon is speed-running obsolescence. |
| |
| ▲ | breuleux an hour ago | parent | next [-] | | If you’re only running models for frontier capabilities, yeah. For tasks where current models are smart enough, running them 100x faster is the most impactful improvement you can make. Consider all the things you could use a model for, but don’t, because the latency is just a bit too high. | |
| ▲ | zxspectrum1982 26 minutes ago | parent | prev | next [-] | | I'd gladly pay for a Claude Opus 4.6 Thinking High in silicon and use it for 1-2 years. It's good enough for many coding tasks. | | |
| ▲ | Gigachad 2 minutes ago | parent [-] | | It costs something like $300,000 for the hardware to run a model of that size. You'd pay that for a single model for 1-2 years? Not even the AI companies can justify that kind of spend which is why they keep extending the expected lifespan on their hardware in the accounting. | | |
| ▲ | zxspectrum1982 a few seconds ago | parent [-] | | I'm expecting the Taalas MSIC verison to cost a fraction of that. Then probably have some kind of cheap subscription to Anthropic for updates (yes, Taalas chips can receive a certain kind of updates: they have a small SRAM). |
|
| |
| ▲ | mdp2021 an hour ago | parent | prev | next [-] | | Compute the cost of producing n of them devices, imagine a fair price based on that, and see if that local, blazing fast card* can be an asset that could be replaced periodically. *(It's local: private files managing firm oriented. It's blazing fast: it can be placed into recursive, intensive local workflows.) | |
| ▲ | try-working 26 minutes ago | parent | prev | next [-] | | obsolescence is the whole point. apple gets to sell a new phone very 6-12 months because of it. i have written about this: "For device makers Packaging models with laptops and smartphones will let application access near free, low latency inference and potentially offer users a better experience with the option of preserving data on-device. This is viable under the condition that tasks that do require larger expert models that run in the cloud can be routed to external models.
A side-effect of local models and what will let Apple cut upgrade cycles from ~4 years (?) down to 12-18 months is specialized hardware to run them. For almost a decade, smartphones have been trying to compete on better cameras. This coming decade will see them selling better GPUs, NPUs, ASICs and whatever other things they'll be calling the inference chips, to drive re-purchase. Every six months will see a better model on new hardware, which will enable better performance in certain applications." https://try.works/role-model-the-case-for-a-model-routing-pr... | | | |
| ▲ | topspin an hour ago | parent | prev | next [-] | | "seems like baking models into silicon is speed-running obsolescence" Now maybe. When models are flying passenger aircraft, other prerogatives will assert themselves. When a 50TB ROM means you can impulse purchase a ChatGPT 6.3 xhigh that runs on batteries, yet more use cases will be apparent. | | |
| ▲ | mdp2021 37 minutes ago | parent [-] | | Well, 50TB ROM Taalas HC1 style would be apparently a 400000b transistor system through a chip sized 2.5 meters on the side... :) | | |
| |
| ▲ | ray_v an hour ago | parent | prev | next [-] | | I could see this making sense when model development start to settle down ... it's going to settle down, right? ... | |
| ▲ | alightsoul an hour ago | parent | prev | next [-] | | Which is exactly what companies and shareholders want to increase sales. | |
| ▲ | amelius an hour ago | parent | prev | next [-] | | Not sure. You can fix the transistors but leave the connections between them open for flexibility, so you only need to change the manufacturing process for the upper masks for every new model. | | |
| ▲ | tsujamin an hour ago | parent | next [-] | | Surely that added flexibility negatively impacts the density/parameter count of the model you could etch? | |
| ▲ | sroussey an hour ago | parent | prev [-] | | Or do a hybrid |
| |
| ▲ | flyinglizard an hour ago | parent | prev [-] | | Look at it the other way: compared to the cost of training a model, the cost of making a custom ASIC is trivial. |
|
|
| ▲ | mrtksn an hour ago | parent | prev | next [-] |
| Isn’t that kind of useless for the stock? It sounds complicated, unlike having number of CPUs go up. It’s like talking about anything else than Megapixels when everyone was convinced that megapixels must go up in certain periods of the smartphone boom. |
|
| ▲ | CircuitSeuss 41 minutes ago | parent | prev | next [-] |
| Apparently Anthropic is moving that way:
https://arstechnica.com/ai/2026/08/anthropic-confirms-plans-... |
| |
| ▲ | mdp2021 15 minutes ago | parent [-] | | Not necessarily: it is relevant to Taalas only if it is a compute-in-memory architecture. The Jalapeño mentioned («Anthropic is not alone in walking this path») in the article is still a classical Von Neumann architecture. And Taalas' idea makes sense in a perspective of scale - producing a large number of cards; "for internal use" (a lower order of items) means a high production cost. |
|
|
| ▲ | giancarlostoro 43 minutes ago | parent | prev | next [-] |
| ASICs is what took over Bitcoin mining, cheaper in all ways, and lasts longer than Nvidia GPUs for inference. |
|
| ▲ | LPisGood an hour ago | parent | prev | next [-] |
| I’m surprised Nvidia hasn’t partnered to make a Claude chip yet. It’s a win/win you can license them out, sell them when they become obsolete, etc. |
|
| ▲ | karmasimida an hour ago | parent | prev | next [-] |
| A model can't be updated, and a chip that is only relevant for 6 months at max? |
| |
| ▲ | anigbrowl 43 minutes ago | parent | next [-] | | Depends what you mean by relevant. If you use AI primarily as a search/knowledge engine, it makes no sense. If it's your capable assistant that has a lot of general knowledge, can do tool calls, and has a big context window, very doable. Indeed, for some kinds of applications involving secure/legal data etc. I can see the consistency of silicon winning out, because it combines performance with immutability and guardrails in hardware. Some chips have write-once PROMs to store password hashes and similar, you could do the same thing with prompt hashing to absolutely force or forbid certain behaviors. A model that can't be updated is also a model that can't be hacked. | |
| ▲ | askvictor an hour ago | parent | prev | next [-] | | People already buy new phones every year, this just creates even more reason to do so | | |
| ▲ | Gigachad a minute ago | parent [-] | | Outside of this website I've never met a person who buys a new phone every year. It's closer to every 3-4 years for most people. |
| |
| ▲ | hamdingers 24 minutes ago | parent | prev [-] | | One of these chips smart enough to take orders at a drive-thru would be relevant for a decade, minimum. |
|
|
| ▲ | alightsoul an hour ago | parent | prev | next [-] |
| Because Openai and anthropic are not hardware companies. They outsource that to Broadcom and AWS' Annapurna labs. |
| |
| ▲ | wmf 38 minutes ago | parent [-] | | OpenAI and Anthropic are both designing ASICs. | | |
| ▲ | alightsoul 22 minutes ago | parent [-] | | So they have decided that putting a small LLM on a phone would backfire because people would have a negative perception of their cloud models. Pretty sure AMD will use these taalas chips in data centers, not phones |
|
|
|
| ▲ | wolttam an hour ago | parent | prev | next [-] |
| It's a terrible moat. You etch the silicon then nobody wants to run it in 6 months because models have advanced that much further. |
| |
| ▲ | anigbrowl 41 minutes ago | parent | next [-] | | This is only true for people who are solely focused on performance. There is absolutely a market for acceptable performance combined with predictability. | |
| ▲ | an hour ago | parent | prev | next [-] | | [deleted] | |
| ▲ | nine_k an hour ago | parent | prev | next [-] | | Not so if it's embedded in something smart enough for its intended purpose. Think vision, spatial reasoning, speech synthesis, even some speech analysis. Think self-driving cars (and drones) that need 10x less power for the brain, and can think at 10x situation per second. | |
| ▲ | twobitshifter 10 minutes ago | parent | prev | next [-] | | OTOH, people get a new iPhone every year and they are ok with it. | | |
| ▲ | nomel 2 minutes ago | parent [-] | | How is that in any way related? A new iPhone gets you a 10% performance improvement of a general purpose CPU. This wouldn't be targeted at generic consumers, probably, for decades. | | |
| |
| ▲ | speed_spread an hour ago | parent | prev [-] | | If a model is good enough today, it's still gonna be good enough in a year. Except you'll be able to serve it 1/100 of the price. Or 100x the speed. |
|
|
| ▲ | bamboozled an hour ago | parent | prev [-] |
| It googles models suck |