| ▲ | bonoboTP an hour ago | |
Exactly. What wasn't really foreseen is that it would come not through some big modeling insight but mainly through massive data scale. People much more imagined some clever general, compact learning algorithm, not quite just gradient descent, but something more intellectually satisfying, and something where you'd feel like "you cracked the mechanism" and you'd see clearly that some critical missing piece had to be invented that unlocked the "understanding" in the model, maybe some kind of fancy Hofstadterian self-referential loop or something. But it just turned out to be data and compute and of course engineering the algorithms to be efficient (which I don't want to discount of course). That's basically the bitter lesson. Academics only reluctantly swallowed that pill and still aren't satisfied with this answer. It's ugly and feels like it shouldn't work because intuition would say there are too many combinations, curse of dimensionality, etc. But it turns out it's just line go up, extrapolate Moore's law and don't worry too much about philosophical-level breakthroughs just count the flops and bits. Ray Kurzweil's scifi extrapolations turned out closer to the truth, whether deservedly or by luck. | ||