44mb of cache really highlights the bottleneck of how raw compute is cheap, but the real limit is memory latency.