Essentially because attention introduced a way to scale un/self-supervised learning to the level of data out there, and learnable inference time 0-shot feature selection. Impressively in a autoregressive, unidirectional manner.