| ▲ | tristor 5 hours ago |
| I wish they were offering 1TB of Unified Memory for the M5 Ultra. I already have an M5 Max MBP w/ 128GB of RAM for running local models, and while there's a /few/ models that I can run in 512GB that I can't run in 128GB that are interesting, where things really shift is at 1TB of memory which allows you run >1T parameter models w/ 4 bit quants reliably. 512GB is just on the edge of "enough", which is maybe the point of maximum frustration considering current memory prices. Personally, I can't justify dropping the dosh for a 512GB M5 Ultra, but I would be able to justify it to myself if I could get 1TB of memory, because it'd guarantee the flexibility with local models I currently am missing. Seems a huge miss to not offer this... for a price. |
|
| ▲ | kamranjon 5 hours ago | parent | next [-] |
| 1tb would likely be ~$20k - given the current >$10k price tag of 256gb. Would you still be considering it at that price? |
| |
| ▲ | jdcasale 5 hours ago | parent | next [-] | | I'd consider a 1tb machine at 20k, but I'm not going to pick up a 256gb one at all. 1TB fits a frontier-ish model in memory without massive quantization, which is a very interesting capability for a non-rack piece of compute. | | |
| ▲ | epolanski 4 hours ago | parent [-] | | But why...? At that point just rent proper GPUs in the cloud, you'd have way more power and pay only what you use for. |
| |
| ▲ | petercooper 5 hours ago | parent | prev | next [-] | | More likely double that, even. I think you'd still see many buyers there. You can spend like $16k alone on a RTX 6000 PRO with a mere 96GB of VRAM now.. | |
| ▲ | tristor 5 hours ago | parent | prev [-] | | I would probably spend up to $30k if I could get 1TB of Unified Memory, because it would allow me a guarantee to run pretty much any local model I want, including >1T parameter models with reasonable quants. I wouldn't be surprised if 512GB is close to $20k when it becomes orderable in October. The justification is less about absolute price and more about price to what it enables. 512GB really doesn't enable much over 128GB for me, but 1TB would massively change things. I have a bit of paranoia/anxiety about AI, but it's not what most people are concerned with. I understand the limits of these tools very well, and still find them extremely useful. What concerns me is that it's going to become difficult to impossible in the future to run local models which have near-SOTA capabilities in a way in which you can exercise full control of the model. I see the writing on the wall, and its more than worth it for me to invest early to ensure my own capabilities. I am very much not a fan of our "you'll own nothing and be happy" directionality for the world, and I am (at least currently) privileged to have the means to slow that decline for my own self. |
|
|
| ▲ | f0cus10 5 hours ago | parent | prev [-] |
| chaining an option? |
| |
| ▲ | DennisP 3 hours ago | parent | next [-] | | They say you can cluster up to four with a shared memory pool, and get three times the inference performance of a single machine. | |
| ▲ | tristor 4 hours ago | parent | prev [-] | | RDMA is buggy and Thunderbolt only delivers 1/10th the throughput of native connectivity. 1TB of Unified Memory w/ 1.2TB/s of bandwidth with marginally ~$30k cost is a different story than 1TB of sorta Unified Memory w/ an effective 120GB/s of bandwidth with a marginally ~$40k cost + all the RDMA bugs. | | |
| ▲ | Lwerewolf 4 hours ago | parent [-] | | You need latency for token parallelism, not bandwidth. Hence actual RDMA that bypasses the software TCP stack (ROCe or whatever). |
|
|