| ▲ | pipsterwo 12 hours ago | |||||||
1/1000 of inference compute is a non-trivial workload at scale. Gartner estimates ~$28B in inference spend for 2026 making this a $28 million dollar per year workload (edit: based on the assumption above) Source: https://www.gartner.com/en/newsroom/press-releases/2026-07-2... | ||||||||
| ▲ | boroboro4 12 hours ago | parent [-] | |||||||
The issue is it’s cpu compute which is underutilized in gpu clusters anyway, so practically it’s not really 1/1000. | ||||||||
| ||||||||