| ▲ | brrrrrm 6 hours ago | |||||||
this is a nice and concise writeup. what's striking to me is that these techniques really have not changed in /years/. sure, precision has become slightly lower, spec decoding acceptance has gotten slightly better and the complexity of parallelism is trickier with mixture of experts. but no new concepts in a very long time! the absolute most impactful improvements for inference comes at architecture design time. I firmly believe everyone who cares about impacting model efficiency should look there | ||||||||
| ▲ | philipkiely 6 hours ago | parent | next [-] | |||||||
I think the biggest net new recent technique is P/D disaggregation. And that spec dec is very different now especially post DSpark/DFlash. But overall yes the fundamentals of LLM performance optimization have been remarkably stable over the last few years. | ||||||||
| ||||||||
| ▲ | Ozzie-D 5 hours ago | parent | prev [-] | |||||||
[flagged] | ||||||||