Remix.run Logo
cmrdporcupine 2 hours ago

People who buy Macs to run local inference are not the target audience for this machine.

It's an AI research workstation for people whose ultimate work goes on production GB300 NVL72 data centre racks (and for large models you link them together.)

It's also about 15x the memory bandwidth of a Mac. For models that fit you'd be looking at hundreds of tokens per second on decode and prefill many times that.

And has CUDA, which is (likely) what your production system will use.

And runs a real server operating system.

Also by the time you spec'd out a Mac with the same total memory capacity and computation you'd also be as expensive. And still not have as many cores, nor have the ConnectX RDMA networking speeds.