| ▲ | ActorNightly 2 hours ago | |||||||
Honest question - why are you so stuck on Macs for local inference? A 4 GPU linux box with 3090s, which are $1500 a piece right now, will blow this thing out of the water. Even 2x3090 rig will run most of the good local models like Gemma4:31b at 100+ tok/sec The VRAM of the GPUs are MUCH faster than the unified ram within Apple Silicon. The only difference is the initial model load, which takes longer from disk to VRAM due to PCIE limitations, but once the model is loaded, GPUs can prefill and and generate tokens way faster than any Apple Silicon. So is your desire to upgrade to mac because you just aren't aware of how to set up a GPU rig, or is it something else? | ||||||||
| ▲ | mft_ 29 minutes ago | parent | next [-] | |||||||
AIUI you can only connect two 3090s at a time with NV link? So you’d only have 48GB of fast combined memory? That’s not tremendously interesting as you’re still limited to the smaller models which, while impressive in their own right, are still IMO too limited to use as your only model. | ||||||||
| ▲ | TruthSHIFT 2 hours ago | parent | prev | next [-] | |||||||
Upgrading RAM is still probably cheaper than spending $6000 on a linux box. You're correct that inference will be much faster on the Linux box. But, the mac's unified memory can load larger models. And as OP mentions, it's still hard to beat cloud pricing at home. | ||||||||
| ▲ | mark_l_watson 43 minutes ago | parent | prev | next [-] | |||||||
Yes, GPUs are much better for dense models. On Macs, MOE models run better, so I agree Macs are more limited and expensive. I have a Linux laptop with a 10GB 1080 GPU, dated, but I should add even more system RAM and try that. | ||||||||
| ▲ | 2 hours ago | parent | prev | next [-] | |||||||
| [deleted] | ||||||||
| ▲ | theshrike79 42 minutes ago | parent | prev [-] | |||||||
A Mac mini running an LLM is quiet A PC with similar capabilities is going to sound like a jet taking off. | ||||||||
| ||||||||