Remix clone Hacker News

new | show | ask | jobs Github

	▲	vjerancrnjak 6 hours ago
		No. There is good signal in IMO gold medal performance. These models actually learn distributed representations of nontrivial search algorithms. A whole field of theorem provingaftwr decades of refinements couldn’t even win a medal yet 8B param models are doing it very well. Attention mechanism, a bruteforce quadratic approach, combined with gradient descent is actually discovering very efficient distributed representations of algorithms. I don’t think they can even be extracted and made into an imperative program.