| ▲ | dist-epoch 5 hours ago |
| > 10 grand for 256GB memory. A NVIDIA RTX 6000, 96 GB at 1.7 TB/s, is 13 grand. This 256 GB at 1.2 TB/s Mac is extremely competitive, it will be sold out everywhere. |
|
| ▲ | blints 5 hours ago | parent | next [-] |
| The relevant comparison isn't one mac studio to one RTX 6000, it's a 24 channel DDR5 system, which also has ~1.2TB/s of memory bandwidth (or more when Xeon 6 compatible 8800mt/s memory becomes widely available), vastly higher prefill due to more CPU horsepower, orders of magnitude faster networking, can hook into GPU accelerators, can be upgraded etc. A baseline 384GB system from eg Puget is ~30K vs ~12K for the 256GB Mac Studio and you do get value for the money. |
| |
| ▲ | ricardobeat 4 hours ago | parent | next [-] | | So 3x more, plus the cost of a GPU (another 10k?). How is that value for money to get slightly better performance? | | |
| ▲ | blints 4 hours ago | parent [-] | | It can be more than "slightly", particularly if the model you're interested in (or will be interested in in 6 months) doesn't fit on the mac studio. You also need to account for eg storing 10TB of random checkpoints, load time when experimenting, and so on. When you start actually needing throughput these are all capability gaps in practical use, not just x% benchmark differences. If you just want to run Qwen 3.8 27B and Deepseek v4 Flash in perpetuity and that's it, there are a lot of solutions that will work and this is a fairly user friendly one. | | |
| ▲ | freehorse 2 hours ago | parent | next [-] | | A 3x price difference means you can get 3 256GB mac studios which you can connect through thunderbolt and with RDMA a total of 768GB ram with compute/memory bandwidth scaling basically linearly (with a small overhead cost). | |
| ▲ | kridsdale1 2 hours ago | parent | prev [-] | | You can find model variants to scale up to whatever capacity you have. Queen has models that just barely fit in 256gb, I ran them… okay… on my Studio. |
|
| |
| ▲ | jauntywundrkind an hour ago | parent | prev | next [-] | | Excellent post. Heck yes. And with MRDIMMs coming, we're going to get another >50% boost in throughput per channel real soon, with a massive uptick in max capacity (4x). It feels like PCIe is a bit of a boat anchor here. There's a SATA->NVMe style transition waiting in the wings to make this all so much better. We really need post-PCIe GPUs. CXL with it's very small low latency flits. This is an "almost certainly not" but I wonder if you could mix PCIe and CXL so you could have the GPU memory expose vmeme as a bunch of CXL.mem pools but still have an otherwise pretty normal GPU. It seems madness that UALink went all in on GPU-to-GPU with no affordances for connecting to host computers. | |
| ▲ | 4 hours ago | parent | prev [-] | | [deleted] |
|
|
| ▲ | petercooper 5 hours ago | parent | prev | next [-] |
| How's the compute side now, I wonder? Because while the Ultras have impressive memory bandwidth for inference, processing prompts still takes a dog's age on my M3 Ultra. I heard the M5 makes some strides forward in this area, though, and the M7 in particular promises to go a lot further. |
| |
| ▲ | dannyw 4 hours ago | parent | next [-] | | M5 is excellent, they’ve finally gotten their own tensor cores. Good for inference; however if you like to train, data format support and effective performance is limited (M5 Pro). Some hardware features are not exposed or extremely slow. You’ll be fine for inference, but pales in comparison to what a RTX 6000 Pro can do for compute/matmuls/training. | |
| ▲ | 4 hours ago | parent | prev [-] | | [deleted] |
|
|
| ▲ | angoragoats 5 hours ago | parent | prev [-] |
| Except the RTX 6000 will run circles around the Mac studio in just about every way. Memory bandwidth is literally the only spec where Apple is competitive, and while high memory bandwidth is necessary for LLMs to perform well, many people strangely don't understand that memory bandwidth alone is not sufficient. |
| |
| ▲ | speedgoose 2 hours ago | parent | next [-] | | Energy efficiency is in another league with the Mac Studio for local workloads. I can run agents using deepseek v4 flash or Qwen 3.8 on my m3 ultra and it will be lukewarm and the fan will eventually start blowing softly. | |
| ▲ | F7F7F7 4 hours ago | parent | prev [-] | | It better because you’ll need a few of them to run some larger models (I’ll be just as vague citing which models). |
|