| ▲ | theplumber 5 hours ago |
| >> 252GB of HBM3e at 7.1TB/s, So it has only 252GB of actual ”AI” memory making it “useless”/toy for actual real world AI workloads(I.e it can’t replace something like opus 5) |
|
| ▲ | theplumber 4 hours ago | parent | next [-] |
| I find it just misleading by advertising 700GB RAM as AI headline. I could plug a 32GB GPU to my 1.5TB ram server and call it “AI station with 1.5TB+ RAM” just so that you find it actually useless compared to the headline |
|
| ▲ | kees99 5 hours ago | parent | prev | next [-] |
| Should work just fine for MoE models where active set fits into 252GB. |
| |
| ▲ | theplumber 4 hours ago | parent [-] | | Can’t you do that already more or less with a Mac Studio with 256 or better 512fb of ram? | | |
| ▲ | fc417fc802 4 hours ago | parent [-] | | Where are you getting 7 TB/s of memory bandwidth? | | |
| ▲ | theplumber 2 hours ago | parent [-] | | Fair point but you get that only for 252GB of ram so it’s not even “totally better” than a Mac Studio ultra with 512GB RAM. It still the old “smaller but faster than a Mac” stuff nvidia sells.Maybe in 1-2 years they will match the Mac 512 but then Apple will release an even bigger Mac |
|
|
|
|
| ▲ | fc417fc802 5 hours ago | parent | prev | next [-] |
| I don't know about "useless" (it seems quite useful to me) but I do feel mislead. It's unified memory in the same way that my current dGPU has unified memory. I guess nvlink-c2c probably (?) doesn't introduce a bottleneck but it's still two distinct arenas with very different performance characteristics. |
|
| ▲ | cmrdporcupine 4 hours ago | parent | prev | next [-] |
| 7.1TB/s of HBM is not "useless" -- that's 30 times the memory bandwidth of my DGX Spark -- and nobody is expecting such a machine to run "Opus 5" on its own. For such large models even datacentre GB300 NVL72 are multiple trays linked together via NVlink etc. This machine has QSFP ports and ConnectX for linking up for larger models. It's a workstation, not a rack. It's for AI researchers. I'd love to have one on (err, under) my desk. What even is this comment? |
| |
| ▲ | theplumber 2 hours ago | parent [-] | | My point is this is “useless” because you can’t run large models. It’s like having a very fast and expensive SSD with a very small capacity. It’s not useless but it is for “real work”. Apple has been providing 512GB ram machines for several years so to say you need a rack to run a large model for personal usage it’s missing the point. You needed a rack of nvidia cards to match an 128GB ram Mac as well a while ago so it’s more of the same. | | |
| ▲ | cmrdporcupine 2 hours ago | parent [-] | | People who buy Macs to run local inference are not the target audience for this machine. It's an AI research workstation for people whose ultimate work goes on production GB300 NVL72 data centre racks (and for large models you link them together.) It's also about 15x the memory bandwidth of a Mac. For models that fit you'd be looking at hundreds of tokens per second on decode and prefill many times that. And has CUDA, which is (likely) what your production system will use. And runs a real server operating system. Also by the time you spec'd out a Mac with the same total memory capacity and computation you'd also be as expensive. And still not have as many cores, nor have the ConnectX RDMA networking speeds. |
|
|
|
| ▲ | LargoLasskhyfv 4 hours ago | parent | prev [-] |
| [dead] |