Remix.run Logo
ethersteeds 4 hours ago

I think a major factor is memory bandwidth. Apple has raised it steadily for each M series generation, and that hasn't plateaued.

Nvidia leads in bandwidth and specialized architecture, but local inference takes off when it's usably fast at much lower cost and power consumption.