| ▲ | jasonjmcghee 4 hours ago | |
M5 prefill is much faster than M4. I've seen benchmarks that show 4-5x faster of M5 Max vs. M4 Max. For local models you're likely using M5 Max, prefill is low thousands of tokens per second, as opposed to, say high hundreds with M4 Max. For larger dense models, some fraction of that, but similar multiple. | ||
| ▲ | smcleod 4 hours ago | parent [-] | |
Yes, I have the M5 Max. But there was no matmul acceleration before the M4 which made things a lot slower. | ||