Remix.run Logo
vadansky 4 hours ago

Sorry for being lazy, but is there a rough breakdown like "You get sonnet level for M5 and Opus for M5 pro, etc.", or is it still speculative. Or put simpler, do you get Opus level for the 256GB M5 Max?

root_axis 3 hours ago | parent | next [-]

For local LLMs with a Mac, rule of thumb is you always want an Ultra (due to memory bandwidth). Even an M1 Ultra is superior to an M6 Pro in this regard.

There are no configurations even close to running something comparable to frontier model variants, they're simply far too large, but something like full precision Qwen 35b or DeepSeek 70b at 50+ t/s is well within available configuration, and potential for plenty of room for large context sizes.

dannyw 30 minutes ago | parent [-]

256GB is enough for DSv4 Flash, expect maybe ~30tg/s, and a lot better profile.

I'm using Flash heavily, and I would describe it as nearly as intelligent as Sonnet-class in agentic coding, but more usable. Less world knowledge of course, and definitely a bit less intelligent; but not _that_ much.

On usability: Takes less handholding, less likely to make unsolicited refactors or whatever, and the writing style is readable.

It's not great at super-long-horizon goals as the Claude 5 models are; but if you have a good harness, you can get around that.

rogerkirkness 4 hours ago | parent | prev | next [-]

Opus is probably ~2T parameter model, so that would probably not run on these. More like Sonnet.

c0rruptbytes 3 hours ago | parent | next [-]

The 512GB could run GLM 5.3 which is Opus level

root_axis 3 hours ago | parent | prev [-]

Sonnet is estimated around 1T, so that is far beyond what's practical as well.

3 hours ago | parent | prev | next [-]
[deleted]
beernet 4 hours ago | parent | prev [-]

[dead]