Remix.run Logo
EagnaIonat 2 hours ago

Your statement is not really true any more.

I have an old M1 Max with 32GB that runs inference models fine, especially those using MLX.

Dedicated graphics cards are faster but not to the point that it matters.

That’s a six year old machine.

My newer machine M5 max 128GB will far outperform your typical 32GB gaming card once the model exceeds memory.

To even get close to that on a dedicated graphics card you are paying upwards of $20K.

That said, models are getting smaller and faster which allows me to run multiple different models with ease.

[edit] Just checked the 5 models I use take up 57GB when all loaded at the same time.