| ▲ | ethersteeds 4 hours ago | |
I think a major factor is memory bandwidth. Apple has raised it steadily for each M series generation, and that hasn't plateaued. Nvidia leads in bandwidth and specialized architecture, but local inference takes off when it's usably fast at much lower cost and power consumption. | ||