Remix.run Logo
pama an hour ago

> Is the energy usage so different between local and cloud inference?

In throughput mode for agentic loads, the energy usage (tok/s/MW) of the new NVidia Vera Rubin hardware is 30x lower than that of the B300 and perhaps 450x lower then the H200 was, which in turn is hundreds of times lower than the inference for single users at home in any non-data-center hardware. It feels like comparing the momentum of an ant to the momentum of an elephant.