| ▲ | jtrn 11 hours ago | |
So basically, we have a clear path to local, viable LLMs in the same ballpark as current frontier LLMs. - Get an M5 Ultra with 512 GB of memory. - A couple of generations of improvements to the base model training. - A couple of generations of improvements to the architecture for running it as efficiently as possible. - A couple of generations of open coding agent harnesses like Pi and OpenCode. And given how fast each generation comes and goes, we are now probably guaranteed superhuman coding assistant that requires 350-ish watts of power, is smaller than a toaster, and can code at 50 TPS. And even if the M5 cost is substantial, is WAY WAY cheaper than what I though would be even possible within a reasonable amount of time. I cant think of anything in history that has improved at this rate... And yea, theres a lot of AI hype, but the amount of people that dont realize how insane this is, suprises me. | ||
| ▲ | nullbio 9 hours ago | parent [-] | |
The first company to release solid open-weight models on hardware LLMs to general consumers will change the world. I don't see a good reason to invest in something like an M5 ultra when it'll be so unbearably slow to generate tokens. Imagine getting stuck with something doing 50 TPS when everyone else is doing 2000-10,000 TPS (10k TPS is what the latest chips hit, I'm not pulling that number out of thin air). | ||