Remix clone Hacker News

new | show | ask | jobs Github

	▲	in-silico 20 hours ago
		Neither of these strike me as particularly groundbreaking. The first idea (as I understand it as retrieving token ids rather than hidden states) is going to really struggle to do useful compositional reasoning and contextual recall. The second idea has been been done a million times, with Linear Attention being maybe the first modern example. Hyena, state-space models, DeltaNet, and LaCT also lie in different regions of the performance-parallelizability spectrum of fixed-size models.