| ▲ | stillpointlab 18 hours ago | |||||||||||||||||||
I admit I would like a faster model - but even though I have faster models available I still go to Fable or GPT-5.6 90% of the time. So there is a gap between a potential preference and a revealed preference. Custom AI for things like facial recognition in cameras has existed for decades, before LLMs were a thing. I don't see that getting replaced. And on-device conversational intelligence might go that route as well, we'll have to wait and see. It's a lot of silicon to dedicated to a static non-changing thing. My money would be on programmable TPU-like things (Apple's NPU kind of stuff). It just seems more flexible to have an array of compute that you can load different models into, so you can update it, etc. | ||||||||||||||||||||
| ▲ | breuleux 11 hours ago | parent | next [-] | |||||||||||||||||||
> even though I have faster models available I still go to Fable or GPT-5.6 90% of the time What about all the things you don't currently use an LLM for? If a specialized chip can run a model 100 times faster, you can suddenly use it for a lot of things at sub-second latency. You can write "make white transparent and add a red outline to x.png" instead of the corresponding imagemagick invocation and perceive little to no latency difference. You can hook it up to your browser and have it yank out all advertisements live, or tell it to highlight anything that might interest you, again, live. There's probably thousands of latent use cases nobody has thought of that would be enabled by a truly fast LLM, even a mediocre one. I don't think an on-device model needs to change much; it's already quite general in its capabilities. | ||||||||||||||||||||
| ||||||||||||||||||||
| ▲ | sdfefcxv 16 hours ago | parent | prev [-] | |||||||||||||||||||
What you think doesn't matter unless its strictly for personal use. Your company will decide what makes economical sense. | ||||||||||||||||||||