Remix.run Logo
xyzzy123 3 hours ago

In my opinion this is one of the areas where GPUs provide a qualitatively different experience than unified memory boxes.

For an EPYC with a 5090 (no layers on CPU) vs an M3 max 128GB, qwen 3.6 27B at 128k context / 7k generation:

                Cold: prefill + decode    Hot (KV cached)
  5090          40s  + 2-3m  = 3-4 min    2-3 min
  M3 Max 128GB  14m  + 8-10m = 22-25 min  8-10 min
This is for dense qwen (which I wouldn't run day to day on the mac) - in reality the mac is quite usable with MoEs but you definitely notice a difference.