| ▲ | SwellJoe 18 hours ago | ||||||||||||||||
"I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now." I don't think it'll take 10-15 years. Gemma 4 31B in the 4-bit QAT is competitive with the frontier of less than three years ago and runs on any high-end 32GB gaming PC GPU or a large-ish Mac. The question is whether the frontier will continue to get better at a rate that allows it to stay ahead of the two curves of availability of consumer hardware big enough to run somewhat larger models and the capability of small models to compete with large ones. When the bottom falls out and GPUs/RAM becomes affordable again, the size of what normal people have on their desk will trend quite a bit larger than today. I think there's a future not too far from now, where a 120B model with really good reasoning and a large context, but limited knowledge (necessitated by being small, you can't fit the world's knowledge in 100 gigabytes), can substitute for a frontier model on almost any task, just by giving it access to web search and documentation for the thing you're trying to do. A 256GB unified memory machine with sufficient memory bandwidth would comfortably run that 120B model. | |||||||||||||||||
| ▲ | tomtheelder 18 hours ago | parent | next [-] | ||||||||||||||||
I think the question is even a bit more nuanced than that. Even if frontier models can maintain a big gap that gap has to actually _matter_. If a local model satisfies my everyday use cases adequately then I may not really care that a frontier model is 5, 10, 100x better at ultra high order reasoning tasks. I think that reality is probably not all that far off for a huge swath of use cases. | |||||||||||||||||
| |||||||||||||||||
| ▲ | satvikpendem 18 hours ago | parent | prev | next [-] | ||||||||||||||||
Hell, Bonsai Labs 27B parameter model can run on phones with their ternary implementation which is quite efficient. Scale that up to frontier model parameters and it's quite likely we can run them on current laptops. | |||||||||||||||||
| ▲ | crubier 12 hours ago | parent | prev [-] | ||||||||||||||||
Came here to say that, my bet is that in 3-4 years you'll be able to run Fable-level of intelligence models on your laptop or maybe even on you phone | |||||||||||||||||
| |||||||||||||||||