Remix.run Logo
rvz 6 hours ago

> M5 Ultra features a massive amount of high-bandwidth unified memory, up to 512GB, and delivers a staggering 1.2TB/s of unified memory bandwidth that is 50 percent higher than M3 Ultra.

Apple never needed to participate in the AI race to zero. Because they were already at the finish line years ago building their own chips that can run large >100B parameter AI models locally.

nasaeclipse 6 hours ago | parent | next [-]

As someone who works in AI now, I have found it pretty amazing that Apple basically didn't do much with AI software, and focused more on the hardware side. I think this is what the future of AI is going to look like, local models run on your mac for your workflow.

It's possible that they're working on their own LLM that's going to work very well on their chips, and possibly outperform anything out there when they do release it.

compounding_it 6 hours ago | parent | next [-]

>local models run on your mac for your workflow.

10 years ago 32GB ram laptops sounded too much. 8 was enough. These days even I would get that much ram since it’s soldered. 64GB is higher end.

In a few years we should see such high end hardware commonplace. Working with a local LLM to get work done is the ideal way to go which has mostly hardware limitation as of now that gets solved in due time.

parineum 6 hours ago | parent [-]

> 10 years ago 32GB ram laptops sounded too much. 8 was enough. These days even I would get that much ram since it’s soldered. 64GB is higher end.

Ten years ago I got 64gb of ram in my laptop, same as I have now. I bought both for business and personal use. System ram capacity hasn't changed much in 10 years.

It makes me curious how old you were 10 years ago.

swiftcoder 5 hours ago | parent [-]

> Ten years ago I got 64gb of ram in my laptop

We were definitely outliers that long ago. I put 64 GB in a MacBook Pro back in 2019, and that was (a) overkill for everything I ever ran on that machine, and (b) stupidly expensive by 2019 standards (albeit almost affordable by 2026 standards)

llm_nerd 5 hours ago | parent | prev | next [-]

>I have found it pretty amazing that Apple basically didn't do much with AI software

The iPhone 15 was almost entirely marketed based upon AI (I would say fraudulently so, advertising features they still haven't delivered), and a huge portion of the OS work was on local AI or AI integration.

And for that matter Apple has been dumping enormous sums into their own AI development. Their failure to have a lot to show for it doesn't void the fact that they tried really, really hard.

It's bizarre how often this "Apple sat on the sidelines and let the AI people fight...so smart!" narrative appears on HN. Apple hasn't gone down the path of spending hundreds of billions on nvidia GPU data centres, but they absolutely tried really hard to matter in AI.

givinguflac 6 hours ago | parent | prev | next [-]

|It's possible that they're working on their own LLM

Yep, Siri AI; they’re doing it in public.

robotresearcher 5 hours ago | parent [-]

‘Apple Foundation Model’

ngvrnd 6 hours ago | parent | prev [-]

nth mover advantage.

Havoc 2 hours ago | parent [-]

Has yet to materialise

LeBit 6 hours ago | parent | prev | next [-]

1.2TB/s is 2/3 the speed of an nVidia 5090.

But you get a generic computer and much more RAM.

And you lose a couple of organs.

bel8 6 hours ago | parent | next [-]

The real downside for me is not having Linux support.

It would take Apple one or two engineers to make Linux life much easier on macs. But Linux is outside their walled garden so it's ignored.

3form 4 hours ago | parent | next [-]

Same here. Sadly I think the voices like ours won't be heard, though, because Apple's looking for someone who's going to buy in on the whole ecosystem, and I think we're not it. Or at least I'm not.

LeBit 6 hours ago | parent | prev [-]

I’m done with macOS.

My Mac Mini is strictly a headless server for llama.cpp.

I use a Linux workstation.

If I were limited to use Mac hardware , I would install Linux in VMware Fusion and work from there.

mhast 5 hours ago | parent | prev | next [-]

It's worth noting that the 5090 (or the RTX Pro 6000 big brother with 92GB VRAM) will run rings around the Mac when it comes to compute.

My old 3090 is typically significantly faster (almost 2x token/s) than my M4 Max 128GB machine, as long as the model fits in the 24GB of VRAM.

In most situations it's a better idea to just buy tokens. But there are definitely cases when that's not an option. And then a machine like the M5 Ultra can allow you to do things locally for a fairly limited budget. And in a simpler package to manage than a machine with multiple GPUs.

snapcaster 6 hours ago | parent | prev | next [-]

is it still effectively 2/3rds? Don't know enough to compare a discrete GPU/CPU setup to something like this where it's more integrated

danielEM 5 hours ago | parent [-]

There is no magic, if the data you compute as atomic chunk don't fit in cache then memory bandwidth R/W limit kicks in and architecture does not matter. On contrary - having multi gpu setup of same price and same memory size with even slower memories may give you effectively much higher bandwidth but at the cost of power consumption.

noodletheworld 6 hours ago | parent | prev [-]

How much memory does that come with?

rbinv 6 hours ago | parent [-]

32 GB GDDR 7

jjice 6 hours ago | parent | prev | next [-]

I was so blown away at all the discourse surrounding "Apple fumbling on models". They should never have been in the model game to begin with. Apple crushes hardware over the last decade and that's a huge advantage today. In the end, massive models have proven to be very strong, but small models have proven to be good enough (especially with the recent Qwen 2.8 27B drop) and that's where I imagine the future will lie for consumers.

chasd00 6 hours ago | parent [-]

I suspected Apple would let everyone else blow all their money then, when the dust settles, deliver a better experience to end users and clean up.

jjice 5 hours ago | parent [-]

Agreed. Apple doesn't innovate anymore, but they're generally pretty good at adapting once other people have.

maherbeg 5 hours ago | parent [-]

I think this is a bit of a crazy statement. Everyone expects Apple to somehow build a category leading product every year. I'd expect something innovative every couple of years

* the iPhone * the iPad * apple watch * airpods * unified memory laptops and computers

Those are all products that either created a category or changed that industry.

jjice 2 hours ago | parent [-]

I think that each of the products you name is the top or near the top of their category, but these weren't creating a category. I don't think my original comment says anything about "changing that industry", so that's a bit of a strawman. They absolutely change the industry they're in. They're just not first to any of those categories that you mentioned (maybe unified memory, I'm not sure).

They weren't the first smart phone, tablet, smart watch, or true wireless earbuds. They did a damn fine job making each of those though. I am typing this on a my work macbook wearing AirPods, and AppleWatch, listening to audio on my iPhone. Apple does a really good job with their products.

Realizing how surrounding by Apple I am...

dgellow 6 hours ago | parent | prev | next [-]

They did participate early on with Apple Intelligence and failed miserably. Really good move to not double down and let the others explore the space first

teekert 6 hours ago | parent | prev [-]

Is there anything comparable that runs Linux, doesn't necessarily look as good, but is perhaps (a lot) cheaper/fixable? Or is this really pretty optimal?

I mean this is not nvidia based right? It's all custom? So we can use it under Asahi perhaps?

I want to get something for my company to run local models, wondering what would be a good option.

datakan 6 hours ago | parent | next [-]

I love linux and would be using it if the ARM support was better. It's just not there and most distros that support ARM do it a little poorly. I just haven't seen anything even remotely comparable to Apple Silicon and unfortunately Linux is struggling very hard to support it.

Marsymars 3 hours ago | parent [-]

It's not quite that ARM support isn't good on Linux, it's that there aren't high-performance ARM chips with strong general-purpose software stacks. Like the Raspberry Pi is very well supported, but otherwise the only upmarket devices are things like Ampere workstations and hyperscaler server chips.

jlokier 5 hours ago | parent | prev | next [-]

You can't run Linux directly on these. Asahi Linux supports up to M2 only.

Linux runs very well in a VM on macOS. There are many good options for this, some free and open source (QEMU, UTM, Lima, Colima), some proprietary (VMware Fusion, Parallels).

But Linux in a VM doesn't get access to the real GPU, so model performance is limited. Those running on the CPU perform well, and those needing the GPU don't.

However, macOS on M-series macs is excellent for local models. (Maybe not as excellent as a box full of the best nVidia GPUs, but still excellent).

So if you're getting Apple hardware, like Linux, and want to run all of it locally, a fine setup for a machine to run local models, with agentic characteristics:

- macOS running one of the many local model runners. I used to use Ollama and Whisper, and now use llama.cpp instead of Ollama. Others use LM Studio, oMLX, etc. Provide HTTP endpoints to access the models.

- Linux in a VM for overall control and orchestration, with standard VM settings, and bridged networking so it appears as its own machine on your network. Also, in here provide a robust shared file server for shared state. Use this VM as your desktop and primary access to the machine, if you like Linux.

- Linux in a VM to launch ephemeral, volatile containers, with the containers using a memory-only tmpfs overlay on top of a read-only Linux filesystem in a VM disk image, with tools in this filesystem. Alternatively, a writable Linux filesystem in a VM disk image, with disk buffering set to use macOS host buffering and discard fsync requests. These settings optimise for container disk performance for data that's only ephemeral which will be deleted soon or on system shutdown. (You can combined both VMs, but need to use two VM disks to get equivalent behaviour, and be careful about VM disk configuration of the two disks.)

- Containers spawned within that second Linux VM can be spawned very quickly and run quickly, so are ideal for LLM agents that need a quick sandbox. These sandboxes generally run faster than a macOS sandbox, despite being on the same machine with VM overhead, because Linux is faster at some things. Teach the LLMs to store files and memories they want to keep in the shared file server.

teekert 5 hours ago | parent | prev | next [-]

I guess, what I mean is: Why are these tiny aluminum boxes so optimal?

I just want my butt ugly repairable beast machine to do the same trick. Why is my ram not unified? I have an iGPU in my server, but it can't access the 64 GB ram (I got last year for 150 euro) directly or something? It's on the CPU right? Why did only Apple go for this architecture? So many questions...

Lunar5227 6 hours ago | parent | prev | next [-]

Strix platform maybe?

mhast 5 hours ago | parent [-]

The PC platforms have anemic memory bandwidth in comparison. Eg, Strix Halo is 256GB/s max. If money is a bigger limiter than performance it can be an option though. As can Nvidia DGX Spark machines. (Also limited to 128GB memory and comparatively low bandwidth, but higher compute than Strix Halo.)

intelkishan 6 hours ago | parent | prev | next [-]

Asahi was stuck at M3 last time I checked it out.

notenlish 5 hours ago | parent | next [-]

Development on m3 is ongoing, m2 is supported

rowanG077 5 hours ago | parent | prev [-]

M2 even.

terminalcommand 6 hours ago | parent | prev [-]

AFAIK, apple does not release drivers open source, asahi is a reverse-engineering endeavour and does not support GPU. For nvidia, there are both proprietary and open-source linux drivers. CUDA and inference works on linux with nvidia. I would recommend checking out this video of Alex Ziskind to shop for a computer to run local LLMs: https://www.youtube.com/watch?v=mevUEQcumzU&t=224s. TL;DR besides Apple he recommends, DGX Spark, Tenstorrent Wormhole N300, AMD Radeon 7900 and NVIDIA RTX 5090.