Remix.run Logo
srcreigh 4 hours ago

Makes you realize how insane the M5 Ultra Mac Studio is. 1.2TB/s bandwidth 512GB memory. Its rated max power draw is just 480W. And it also has amazing M-series CPUs. It costs less than just one of these GPUs which each take 700W to run.

segmondy 3 hours ago | parent | next [-]

If you want to see how impressive Nvidia is, serve 32 concurrent request on it and compare the same with the mac.

No comparison.

None.

srcreigh 2 hours ago | parent | next [-]

There is comparison actually. I spent all day researching this a few days ago.

Memory-wise, the RTX PRO 6000 can barely hold two 1M context Qwen 3.8 27B models at 8 bit quantization at the same time. The 512GB M5 Ultra Mac Studio could hold around 14.

At such high concurrency, batch performance is usually limited more by memory bandwidth than compute. The RTX PRO 6000's memory bandwidth is just 50% faster than the M5 Ultra.

So yeah, I think if we are talking about many short context requests, sure. But if you are chewing through a backlog of coding tasks with Qwen overnight, they might actually be comparable.

Im sure in Nov when the M5 Ultra comes out we'll see a lot of interesting benchmarks.

2 hours ago | parent | prev [-]
[deleted]
qeternity 3 hours ago | parent | prev | next [-]

*TB/s

teaearlgraycold 4 hours ago | parent | prev [-]

These GPUs are extremely inflated in price because Nvidia effectively has a monopoly on hardware that is used to train models. Apple Silicon tends to have good inference software available but as soon as you want to train even a YOLO model bits and pieces fall back to software implementations. Try to train an LLM and it'll get even worse.

The M3 Ultra's GPU performance is around a 4070 Ti. The M5 Ultra more like a 5080. They're both amazing deals compared to Nvidia for local inference because of their massive pool of high bandwidth memory. But a single RTX PRO 6000 should be 2 or 3x the compute of an M5 Ultra.

segmondy 3 hours ago | parent [-]

I can't stand Nvidia, but these GPUs are not inflated because of monopoly. They are expensive because demand exceeds supply. That's it. The world wants compute and we don't have enough of it!

teaearlgraycold 3 hours ago | parent [-]

I said inflated. They would always be expensive. But they charge more than their competitors per flop and per byte because they’re the only manufacturer that can run CUDA. If they lost that moat their prices would go down a notch.