Remix.run Logo
cmrdporcupine 4 hours ago

7.1TB/s of HBM is not "useless" -- that's 30 times the memory bandwidth of my DGX Spark -- and nobody is expecting such a machine to run "Opus 5" on its own. For such large models even datacentre GB300 NVL72 are multiple trays linked together via NVlink etc. This machine has QSFP ports and ConnectX for linking up for larger models.

It's a workstation, not a rack. It's for AI researchers. I'd love to have one on (err, under) my desk.

What even is this comment?

theplumber 2 hours ago | parent [-]

My point is this is “useless” because you can’t run large models. It’s like having a very fast and expensive SSD with a very small capacity.

It’s not useless but it is for “real work”. Apple has been providing 512GB ram machines for several years so to say you need a rack to run a large model for personal usage it’s missing the point.

You needed a rack of nvidia cards to match an 128GB ram Mac as well a while ago so it’s more of the same.

cmrdporcupine 2 hours ago | parent [-]

People who buy Macs to run local inference are not the target audience for this machine.

It's an AI research workstation for people whose ultimate work goes on production GB300 NVL72 data centre racks (and for large models you link them together.)

It's also about 15x the memory bandwidth of a Mac. For models that fit you'd be looking at hundreds of tokens per second on decode and prefill many times that.

And has CUDA, which is (likely) what your production system will use.

And runs a real server operating system.

Also by the time you spec'd out a Mac with the same total memory capacity and computation you'd also be as expensive. And still not have as many cores, nor have the ConnectX RDMA networking speeds.