Remix.run Logo
blints 5 hours ago

10 grand for 256GB memory. Likely double that for 512GB, but won't be available or finalized until October. Thunderbolt 5 is highest bandwidth external IO available at 120Gb/s. 1.2TB/s claimed max internal memory bandwidth.

Not exactly "future proof" for >1T parameter models but good for targeting specific lower-parameter models, or if you can rely on pipeline parallelism and run a cluster.

root_axis 3 hours ago | parent | next [-]

There is no "future proof" for >1T param models, there is no present or future where you can run a model that size on consumer hardware.

bewareofscams 3 hours ago | parent | next [-]

6 Mac Studios for 100k USD. Consider it a rule of thumb now - 100k to run 1T params, scales linearly.

blints 3 hours ago | parent | prev [-]

I don't know that "consumer hardware" is a useful distinction anymore, it's just "what's your budget and what's your speed requirement".

kridsdale1 2 hours ago | parent [-]

Consumer hardware means:

- 120v input plug

- not rack-mounted

- has a video out port

iAMkenough 2 hours ago | parent [-]

Ah, a MacBook Professional is Consumer /s

kridsdale1 an hour ago | parent | next [-]

Yes?

29 minutes ago | parent | prev [-]
[deleted]
BugsJustFindMe 5 hours ago | parent | prev | next [-]

> Not exactly "future proof"

Computers are never "future proof".

chmod775 42 minutes ago | parent | next [-]

They do however last a much longer time now.

A 8 years old graphics card can still play modern games. Ten years ago playing a modern game on hardware that old would've been unthinkable.

And with how the market is right now, we'll be stuck on the current "reference level" of hardware for a while longer.

kridsdale1 2 hours ago | parent | prev | next [-]

I’d say something like a PS4 is future proof. $350 got like 10 years of modern software.

mannanj 4 hours ago | parent | prev | next [-]

About 20 years ago my dad bought me a $5k computer, it was future proof for about "5 years" before we had to upgrade its internal parts (more memory, new graphics card).

It was future proof but not really because it struggled a lot in its final years.

i5heu 3 hours ago | parent | next [-]

My 6 year old mediocre gaming PC is still a mediocre gaming PC.

I expect it to stay a mediocre gaming PC for the next 3 years maybe 5 years.

16GB RAM RTX 3060 TI (8 GB VRAM)

sschueller 2 hours ago | parent | prev [-]

He had that option,l. However buying apple hardware todqy you are stuck and need to replace all of it even if for example the CPU is plenty fast but you need more memory or GPU power.

moomoo11 an hour ago | parent | prev | next [-]

i'd say it depends (as with everything)

i'm still using an old i7 3770k @ 4.8ghz with 16gb (ddr3) ram running linux for random tasks like executing tests. obviously, power consumption is higher.

my main machine is a MBP M1 Max which i use for everything. i also have my main linux desktop workstation that has a 5950x with 128gb (ddr4) ram.

i'll probably get 2-5 years out of my MBP, and my AMD workstation will probably be good for another 5-10 years.

i'm not a gamer, but i have a 3080. i'm sure my 5950x will still be good for gaming in 10 years if paired with a modern GPU.

mschuster91 4 hours ago | parent | prev [-]

> Computers are never "future proof".

Upgradeable components however could go a loooong stretch towards that goal. It can't be that hard to follow a common form factor for at least the housing across two or three generations to allow a reuse of everything but the main PCB.

jve 4 hours ago | parent | next [-]

Think it would have same memory bandwidth if the RAM was upgradeable?

Would be nice if someone knowledgeable about electrical engineering and manufacturing processes could lay out some valid reasons for manufacturers to integrate RAM onto the motherboard.

https://news.ycombinator.com/item?id=49041256#49082206

enragedcacti 3 hours ago | parent [-]

It isn't integrated into the motherboard, it's integrated onto the same package as the CPU/GPU which allows for better signal integrity and higher speeds. You can get somewhat close to the same speeds while modular with tech like LPCAMM2, but there are some pretty difficult challenges to overcome to close the gap completely. As an example, the Framework Laptop 13 Pro CPUs support up to 9600MT/s (same as M5), and Micron sells LPCAMM2 modules that can run at 8533MT/s, but the 13 Pro only officially supports 7467MT/s.

_kush 4 hours ago | parent | prev | next [-]

If it was upgradable, then yes, spending more on top of it every year would make it future proof, but that's not the point. It's that spending 10 grand doesn't get you a future proof computer today.

mhast 4 hours ago | parent | prev | next [-]

The main PCB is pretty much everything that has value. The rest is a heatsink, case and PSU.

fearmerchant 4 hours ago | parent | prev [-]

The way the Apple M-series does ram that might be difficult to pull off.

mschuster91 4 hours ago | parent [-]

> The way the Apple M-series does ram that might be difficult to pull off.

Well it might be an idea to keep the layout of the mainboard and connectors the same.

That way, instead of having to upgrade the whole machine, all it would need is a new mainboard. Framework for example managed to pull that off, and in mobile at that, where constraints are much worse than for a desktop computer.

isgb 4 hours ago | parent | next [-]

> Framework for example managed to pull that off, and in mobile at that, where constraints are much worse than for a desktop computer.

It's not the same thing though. On the M-series, CPU and GPU share a unified memory architecture and ram is much more tightly coupled to get it to go faster. A closer example would be the Framework desktop, actually, where memory is also soldered in for the same reason.

roughly 3 hours ago | parent | prev [-]

That actually would be interesting - yes, the computer itself is basically a PCB, but it’s also wrapped in a couple pounds of aluminum, a power supply, cooling fans, and a few other things that don’t need to be consumables. Upgrade a mini or a studio by swapping the new board into the old case - yeah, you’re not saving much money, but you also don’t need to throw out the entire rest of the case, and you can ship the main board in the space of a couple CD cases.

It’s a very non-Apple thing to do, but it’d be pretty awesome if they did.

SwellJoe 3 hours ago | parent | prev | next [-]

That's in the same ballpark as two 128GB AI machines like the Asus GX10 or DGX Spark or Strix Halo. And, it seems very likely to perform better than either of those for inference. And, 256GB brings some pretty good models into play.

But, that doesn't make it a good deal. It just means the Apple tax doesn't apply when stacked up against AI machines and with memory prices being so out of whack. I'm still planning to wait until the RAMpocalypse ends before I buy any more hardware.

dist-epoch 5 hours ago | parent | prev | next [-]

> 10 grand for 256GB memory.

A NVIDIA RTX 6000, 96 GB at 1.7 TB/s, is 13 grand.

This 256 GB at 1.2 TB/s Mac is extremely competitive, it will be sold out everywhere.

blints 5 hours ago | parent | next [-]

The relevant comparison isn't one mac studio to one RTX 6000, it's a 24 channel DDR5 system, which also has ~1.2TB/s of memory bandwidth (or more when Xeon 6 compatible 8800mt/s memory becomes widely available), vastly higher prefill due to more CPU horsepower, orders of magnitude faster networking, can hook into GPU accelerators, can be upgraded etc. A baseline 384GB system from eg Puget is ~30K vs ~12K for the 256GB Mac Studio and you do get value for the money.

ricardobeat 4 hours ago | parent | next [-]

So 3x more, plus the cost of a GPU (another 10k?). How is that value for money to get slightly better performance?

blints 4 hours ago | parent [-]

It can be more than "slightly", particularly if the model you're interested in (or will be interested in in 6 months) doesn't fit on the mac studio. You also need to account for eg storing 10TB of random checkpoints, load time when experimenting, and so on. When you start actually needing throughput these are all capability gaps in practical use, not just x% benchmark differences.

If you just want to run Qwen 3.8 27B and Deepseek v4 Flash in perpetuity and that's it, there are a lot of solutions that will work and this is a fairly user friendly one.

freehorse 2 hours ago | parent | next [-]

A 3x price difference means you can get 3 256GB mac studios which you can connect through thunderbolt and with RDMA a total of 768GB ram with compute/memory bandwidth scaling basically linearly (with a small overhead cost).

kridsdale1 2 hours ago | parent | prev [-]

You can find model variants to scale up to whatever capacity you have. Queen has models that just barely fit in 256gb, I ran them… okay… on my Studio.

jauntywundrkind an hour ago | parent | prev | next [-]

Excellent post. Heck yes. And with MRDIMMs coming, we're going to get another >50% boost in throughput per channel real soon, with a massive uptick in max capacity (4x).

It feels like PCIe is a bit of a boat anchor here. There's a SATA->NVMe style transition waiting in the wings to make this all so much better. We really need post-PCIe GPUs. CXL with it's very small low latency flits. This is an "almost certainly not" but I wonder if you could mix PCIe and CXL so you could have the GPU memory expose vmeme as a bunch of CXL.mem pools but still have an otherwise pretty normal GPU. It seems madness that UALink went all in on GPU-to-GPU with no affordances for connecting to host computers.

4 hours ago | parent | prev [-]
[deleted]
petercooper 5 hours ago | parent | prev | next [-]

How's the compute side now, I wonder? Because while the Ultras have impressive memory bandwidth for inference, processing prompts still takes a dog's age on my M3 Ultra. I heard the M5 makes some strides forward in this area, though, and the M7 in particular promises to go a lot further.

dannyw 4 hours ago | parent | next [-]

M5 is excellent, they’ve finally gotten their own tensor cores.

Good for inference; however if you like to train, data format support and effective performance is limited (M5 Pro). Some hardware features are not exposed or extremely slow.

You’ll be fine for inference, but pales in comparison to what a RTX 6000 Pro can do for compute/matmuls/training.

4 hours ago | parent | prev [-]
[deleted]
angoragoats 5 hours ago | parent | prev [-]

Except the RTX 6000 will run circles around the Mac studio in just about every way. Memory bandwidth is literally the only spec where Apple is competitive, and while high memory bandwidth is necessary for LLMs to perform well, many people strangely don't understand that memory bandwidth alone is not sufficient.

speedgoose 2 hours ago | parent | next [-]

Energy efficiency is in another league with the Mac Studio for local workloads.

I can run agents using deepseek v4 flash or Qwen 3.8 on my m3 ultra and it will be lukewarm and the fan will eventually start blowing softly.

F7F7F7 4 hours ago | parent | prev [-]

It better because you’ll need a few of them to run some larger models (I’ll be just as vague citing which models).

cma 5 hours ago | parent | prev | next [-]

3 of those thunderbolt 5 ports, so you can do a fully connected 4 machine cluster topology.

blints 5 hours ago | parent [-]

It's unclear to me how bandwidth scales with multiple connections. Many-to-many does not seem ideal. Daisy chaining would be fine for straight pipeline work. There doesn't seem to be an equivalent of a ethernet switch for thunderbolt 5 though.

kridsdale1 2 hours ago | parent [-]

Apple put on the announce page that 4 Studios in this config can run inference at 3x the speed as 1 Studio.

seanmcdirmid 3 hours ago | parent | prev [-]

10 grand for 256GB new Ultra sounds too cheap in today’s crazy DRAM market, it feels too good to be true.