Remix.run Logo
theplumber 2 hours ago

My point is this is “useless” because you can’t run large models. It’s like having a very fast and expensive SSD with a very small capacity.

It’s not useless but it is for “real work”. Apple has been providing 512GB ram machines for several years so to say you need a rack to run a large model for personal usage it’s missing the point.

You needed a rack of nvidia cards to match an 128GB ram Mac as well a while ago so it’s more of the same.

cmrdporcupine 2 hours ago | parent [-]

People who buy Macs to run local inference are not the target audience for this machine.

It's an AI research workstation for people whose ultimate work goes on production GB300 NVL72 data centre racks (and for large models you link them together.)

It's also about 15x the memory bandwidth of a Mac. For models that fit you'd be looking at hundreds of tokens per second on decode and prefill many times that.

And has CUDA, which is (likely) what your production system will use.

And runs a real server operating system.

Also by the time you spec'd out a Mac with the same total memory capacity and computation you'd also be as expensive. And still not have as many cores, nor have the ConnectX RDMA networking speeds.