Remix.run Logo
vmware508 an hour ago

Apple will release M7 MacBook Pros / Mac Minis next year, and they will be able to run free LLMs locally at native speed. All software developer notebooks will be replaced to run local models, saving a lot by cancelling Claude Code subscriptions. Developers win. Apple stocks will be rocketing. Everything else will go down. You're welcome.

schleck8 an hour ago | parent | next [-]

You'd need the 256 gb memory model which will be expensive because apple has trouble getting capacity (got turned down by cxmt). And even then you can only run a 2 bit quant which is noticeably worse than 8 bit

Gecko4072 an hour ago | parent | prev | next [-]

They will cost an insane amount as well. Maybe less than subscriptions or tokens. But running massive models on laptops with batteries and poor cooling doesn’t make much sense.

LeBit 38 minutes ago | parent [-]

Until hiding PII from the cloud LLM is a resolved issue, running local LLMs will remain a necessity.

There are workplaces that refuse to use LLMs because they fear the devs will expose sensitive data without care.

scotty79 3 minutes ago | parent | prev | next [-]

I don't know why you'd want to burden your laptop with a large model. But I can totally see a new "developer workstation" product that's just a semi-large box that's optimized for running frontier open weights models for one to few users.

toasty228 7 minutes ago | parent | prev | next [-]

Sure buddy, all you'll end up with is a $10k machine that run gimped models at like 30tok/s for about 5m before the fan kicks in and it starts to sound like a turboprop, while offering maybe 30% of the context size of hosted models.

Flavius an hour ago | parent | prev | next [-]

> run free LLMs locally at native speed

This reads like a hallucination. What does native speed even mean?

kyxsc 37 minutes ago | parent | next [-]

for example, models running at like 100-150 tokens/second (or faster!) vs 15 t/s

(fable/sol are ~60 t/s, and OpenAI just announced their Cerebras partnership(?) for "ultrafast" mode of 750 t/s)

models aren't able to run that fast right now on our consumer/prosumer hardware. M5 Max for example has a memory bandwidth of 600 GB/s. a 5090 has 3x that, so running the same model on a 5090 is that much faster (provided the model is within 30GB).

running a bigger model on an M5 Ultra is still much slower than running it on a Blackwell chip with sufficient vram, CUDA being a major difference. if apple can bridge this gap, interesting things will happen... and just imagine if M7 Ultra has comparable speeds to Blackwell (or even Rubin)!

toasty228 4 minutes ago | parent [-]

Meanwhile the GB300 used by hosted llms:

GPU Memory Bandwidth: 7.1 TB/s Interconnect Bandwidth: 900 GB/s bidirectional

https://pi3g.com/nvidia-gb300-specifications-including-memor...

If you think M7 will hit even 15% of these speeds you're very optimistic.

lmpdev 35 minutes ago | parent | prev [-]

I assume they mean same t/sec as a SOTA cloud model

re-thc 18 minutes ago | parent | prev [-]

> Apple will release M7 MacBook Pros / Mac Minis next year

The latest on Apple is TSMC is stuck on the next iPhone due to lack of RAM. Good luck getting any Macs. Memory shortage is getting worse.