Remix.run Logo
_doctor_love 4 days ago

Easy prediction: LLMs will get shrunk down further and further until GenAI is just something that ships on a chip as part of your hardware. In the future it will seem quaint that we needed a network connection to talk to our LLM.

Adoption of open-source models to my mind is a similar step in that direction. In all cases, the goal is to become untethered from a mercurial vendor.

honr 4 days ago | parent | next [-]

Like taalas.com (very recently acquired by AMD), or cerebras.ai (whole wafer is a chip)? As you said, I also think that is one of the main direction many companies (and academia) is moving to.

hparadiz 4 days ago | parent | prev | next [-]

We're gonna start baking in models like TTS with thousands of voices available in any language as a chip on device. They just need to hit 99% accuracy and then it's a done deal.

wnmurphy 4 days ago | parent | prev | next [-]

Yeah, I'm looking forward to this actually.

https://chatjimmy.ai/ blew my mind at how fast etched model weights can be.

For on-device LLMs, there's a point of diminishing returns, meaning you don't need to have the latest frontier model for most operations.

hparadiz 4 days ago | parent [-]

I currently have a small TTS model running in the background on my machine through which my agent(s) speak to me as they work. If that can be baked into an ASIC along with a few thousand voices in every major language then it should just be a utility chip on your mobo for anything that needs it. And yes, I too, am looking forward to it.

andriy_koval 4 days ago | parent | prev [-]

for every small GenAI model there will be larger model or cluster of models which are smarter than small model