Remix.run Logo
mark_l_watson 2 hours ago

I run a lot of local models (I am always experimenting) on my 32G M2-Pro MacMini - I would love to upgrade.

The financial aspects don’t work however: I can learn and experiment with what I have for local models, and I pay as I go on FireWorks.ai for open model inferencing and no matter how much I use this service my monthly bill is between $10 and $40 and much faster than any reasonable home rig.

Hybrid ‘small local’ and buying inference is the way I choose.

ActorNightly an hour ago | parent [-]

Honest question - why are you so stuck on Macs for local inference?

A 4 GPU linux box with 3090s, which are $1500 a piece right now, will blow this thing out of the water. Even 2x3090 rig will run most of the good local models like Gemma4:31b at 100+ tok/sec

The VRAM of the GPUs are MUCH faster than the unified ram within Apple Silicon. The only difference is the initial model load, which takes longer from disk to VRAM due to PCIE limitations, but once the model is loaded, GPUs can prefill and and generate tokens way faster than any Apple Silicon.

So is your desire to upgrade to mac because you just aren't aware of how to set up a GPU rig, or is it something else?

TruthSHIFT an hour ago | parent | next [-]

Upgrading RAM is still probably cheaper than spending $6000 on a linux box. You're correct that inference will be much faster on the Linux box. But, the mac's unified memory can load larger models. And as OP mentions, it's still hard to beat cloud pricing at home.

an hour ago | parent | prev [-]
[deleted]