| ▲ | LunaSea a day ago | |||||||||||||||||||||||||
Why? GPUs are replaced every 3 to 5 years. This is going to be an ongoing operational cost forever. It will probably increase more if larger models require bigger VRAM sizes. | ||||||||||||||||||||||||||
| ▲ | mcbuilder a day ago | parent | next [-] | |||||||||||||||||||||||||
We have probably hit a limit to scaling LLMs through raw parameter count alone, at least we're not seeing the exponential pace. I personally think we'll end up with a nice sigmoid curve plateauing in the sub 10T parameter regime. The amount of tokens processed (in inference) is increasing exponentially though (I've been following open router usage stats for years and it's always been exponential). We will of course make technological advances in hardware efficiency, and model parameter efficiency, but I think a much more plausible future is that VRAM needed for loading and serving individual models will slow down or even stop. We will need more chips, and more power, as demand continues to grow of course, but the operational lifetime of GPUs today will be a lot longer than the SoTA cards from 5 years ago. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||
| ▲ | ody4242 a day ago | parent | prev | next [-] | |||||||||||||||||||||||||
They are building new datacenters for the AI demand, so around half of this CAPEX is not for the GPU-s, and those will not be replaced every 3-5 years. Also, TPUv2 was introduced in 2018, and still not completely retired in all regions, from accounting pov, it has been written down to 0, but they are still working. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||
| ▲ | benoau a day ago | parent | prev [-] | |||||||||||||||||||||||||
That cost has always been there and allowed for their lucrative margins. It's the upfront cost of building/populating their datacenters (many more than before) that is eating those margins. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||