Remix.run Logo
iLoveOncall 3 hours ago

> We propose intelligence per watt (IPW), task accuracy per unit of power

Stupid metric. It's not because a model is better performing that it necessarily requires more energy or compute.

utopiah 3 hours ago | parent | next [-]

I don't think that's what they are saying. In fact if they did the metric would be pointless. Rather they are saying by estimating that value on different architectures, one can find more efficient ones. They use open model to be able to remove unknowns. They aren't advocating for one model or another, only more efficient architectures.

the8472 an hour ago | parent [-]

Intelligence per Joule would be more appropriate in many cases. If a model can do the same work but takes 10 times as long as a bigger one that can still be useful (e.g. due to memory constraints), but at the same wattage it burns 10 times the energy. Even more so on mobile devices.

frumiousirc 21 minutes ago | parent [-]

They also define and measure an "IPJ" as well as "IPW"

> the NVIDIA B200 achieves 1.6× to 2.3× higher intelligence per joule than the APPLE M4 MAX across QWEN 3 and GPT-OSS model variants

The B200 = "cloud", M4 = "local".

So "cloud" does even better in energy than it does in power compared to "local". Or, to flip it, "local" is both slower and more expensive than "cloud".

_diyar 3 hours ago | parent | prev [-]

> We propose miles per hour (MPH), distance travelled per unit of time

Stupid metric. It‘s not because you spend more time that you travel farther.

s/