Remix.run Logo
onlyrealcuzzo 42 minutes ago

What makes this insanely hard to predict is that the compute needed for the same quality output has roughly gone down 90% every 18 months for ~5 years.

1) We don't know how long that trend will continue, but you do know where to look for when it may end (if smaller sized models continue to compress the knowledge effectively of larger models).

2) We don't know when the appetite for higher cost models might go down and by how much if smaller models get "good enough" and price becomes far more important.

It is entirely possible that 5 years from now, there's >100x LLM inference going on - but demand for AI chips is only 2x or less.

It is also entirely possible that at some size - LLMs pick up some emergent capability that doesn't scale well to smaller sizes - and that there's an incredible boost to demand to get that capability.

It's just very hard to predict.

mattnewton a minute ago | parent [-]

I think efficiency is unlikely to result in lower demand for compute, instead more useful compute per watt increase the value of that compute and we are not going to run out of economically useful things to do with it anytime soon.