| ▲ | ActorNightly an hour ago | |
Don't worry, there are exactly zero Apple chips that are good for any sort of local inference right now. Unified ram is fast for ram, but dogshit slow compared to actual dedicated VRAM on graphics cards. | ||
| ▲ | EagnaIonat an hour ago | parent [-] | |
Your statement is not really true any more. I have an old M1 Max with 32GB that runs inference models fine, especially those using MLX. Dedicated graphics cards are faster but not to the point that it matters. That’s a six year old machine. My newer machine M5 max 128GB will far outperform your typical 32GB gaming card once the model exceeds memory. To even get close to that on a dedicated graphics card you are paying upwards of $20K. That said, models are getting smaller and faster which allows me to run multiple different models with ease. [edit] Just checked the 5 models I use take up 57GB when all loaded at the same time. | ||