Remix.run Logo
blints 5 hours ago

The relevant comparison isn't one mac studio to one RTX 6000, it's a 24 channel DDR5 system, which also has ~1.2TB/s of memory bandwidth (or more when Xeon 6 compatible 8800mt/s memory becomes widely available), vastly higher prefill due to more CPU horsepower, orders of magnitude faster networking, can hook into GPU accelerators, can be upgraded etc. A baseline 384GB system from eg Puget is ~30K vs ~12K for the 256GB Mac Studio and you do get value for the money.

ricardobeat 4 hours ago | parent | next [-]

So 3x more, plus the cost of a GPU (another 10k?). How is that value for money to get slightly better performance?

blints 4 hours ago | parent [-]

It can be more than "slightly", particularly if the model you're interested in (or will be interested in in 6 months) doesn't fit on the mac studio. You also need to account for eg storing 10TB of random checkpoints, load time when experimenting, and so on. When you start actually needing throughput these are all capability gaps in practical use, not just x% benchmark differences.

If you just want to run Qwen 3.8 27B and Deepseek v4 Flash in perpetuity and that's it, there are a lot of solutions that will work and this is a fairly user friendly one.

freehorse 2 hours ago | parent | next [-]

A 3x price difference means you can get 3 256GB mac studios which you can connect through thunderbolt and with RDMA a total of 768GB ram with compute/memory bandwidth scaling basically linearly (with a small overhead cost).

kridsdale1 2 hours ago | parent | prev [-]

You can find model variants to scale up to whatever capacity you have. Queen has models that just barely fit in 256gb, I ran them… okay… on my Studio.

jauntywundrkind an hour ago | parent | prev | next [-]

Excellent post. Heck yes. And with MRDIMMs coming, we're going to get another >50% boost in throughput per channel real soon, with a massive uptick in max capacity (4x).

It feels like PCIe is a bit of a boat anchor here. There's a SATA->NVMe style transition waiting in the wings to make this all so much better. We really need post-PCIe GPUs. CXL with it's very small low latency flits. This is an "almost certainly not" but I wonder if you could mix PCIe and CXL so you could have the GPU memory expose vmeme as a bunch of CXL.mem pools but still have an otherwise pretty normal GPU. It seems madness that UALink went all in on GPU-to-GPU with no affordances for connecting to host computers.

4 hours ago | parent | prev [-]
[deleted]