| ▲ | the8472 an hour ago | |
Intelligence per Joule would be more appropriate in many cases. If a model can do the same work but takes 10 times as long as a bigger one that can still be useful (e.g. due to memory constraints), but at the same wattage it burns 10 times the energy. Even more so on mobile devices. | ||
| ▲ | frumiousirc 24 minutes ago | parent [-] | |
They also define and measure an "IPJ" as well as "IPW" > the NVIDIA B200 achieves 1.6× to 2.3× higher intelligence per joule than the APPLE M4 MAX across QWEN 3 and GPT-OSS model variants The B200 = "cloud", M4 = "local". So "cloud" does even better in energy than it does in power compared to "local". Or, to flip it, "local" is both slower and more expensive than "cloud". | ||