Remix.run Logo
GodelNumbering 3 hours ago

1.2 TB/s bandwidth of M5 Ultra comes from two dies of M5 Max (each 614 GB/s) connected together using 4.4 TB/s inter-die fabric.

For a non-quantized Deepseek V4 flash on an ultra, I would estimate about 1000+ tokens per second prefill and 50+ tokens per second on generation. This is actually quite usable and near parity to cloud.

They mention "adds the GPU Neural Accelerators." which, if exploitable for LLM loads, would probably help the prefill a lot

hbbio 3 hours ago | parent | next [-]

Yes, and they specifically mention "Up to 10.7x faster LLM prompt processing in LM Studio" which is probably using the neural accelerator for prefill.

dakolli 2 hours ago | parent | prev | next [-]

How much would is the cost for that machine though, I'm pretty sure I could just buy tokens from a provider and never run out of money for 10 years, and get far better quality output because inference is being served by professionals on far better hardware and this machine would be obsolete long before that as well. Hosting local seems like a possibly the dumbest thing you could possibly do from an economics perspective. And don't hit me with the privacy argument because everyone saying they care about privacy uses fucking gmail, whatsapp and instagram all day long.

rxyz 26 minutes ago | parent | next [-]

You can’t run multiple workers 24/7 for 100 bucks a month

kridsdale1 2 hours ago | parent | prev [-]

Ok sure. But this machine you can haul in an Ice Road Truck to the North Pole and do inference there in an off grid shack. Good luck talking to cloud AI there.

dakolli an hour ago | parent [-]

Yeah, that's who Apple is making these for fucking Santa and his Elves lmfao

coder543 an hour ago | parent | prev [-]

[dead]