| ▲ | segmondy an hour ago | |
If you want to see how impressive Nvidia is, serve 32 concurrent request on it and compare the same with the mac. No comparison. None. | ||
| ▲ | srcreigh 13 minutes ago | parent | next [-] | |
There is comparison actually. I spent all day researching this a few days ago. Memory-wise, the RTX PRO 6000 can barely hold two 1M context Qwen 3.8 27B models at 8 bit quantization at the same time. The 512GB M5 Ultra Mac Studio could hold around 14. At such high concurrency, batch performance is usually limited more by memory bandwidth than compute. The RTX PRO 6000's memory bandwidth is just 50% faster than the M5 Ultra. So yeah, I think if we are talking about many short context requests, sure. But if you are chewing through a backlog of coding tasks with Qwen overnight, they might actually be comparable. Im sure in Nov when the M5 Ultra comes out we'll see a lot of interesting benchmarks. | ||
| ▲ | 26 minutes ago | parent | prev [-] | |
| [deleted] | ||