Remix.run Logo
cornholio 20 hours ago

The field now favors into the view that symbolic manipulation is not the mechanism of general intelligence, but rather an emergent byproduct of learning. So the fact that a connectionist machine (neural network) got so good at symbolic manipulation actually supports the view that we are closing the gap to general intelligence. Through the rote work, the machine really internalizes those rules and the symbolic manipulation capabilities are emergent, just like we humans do it.

What still confuses people is the insane inefficiency of deep learning, and that those emergent capabilities require such an immense training corpus compared to the only other architecture that we know of.

But this already is an optimization problem. If the machine gets super human at symbolic reasoning, and at the same time, can solve the symbol grounding problem to real world data and sensors, what prevents you from saying it thinks? Can it not solve real world problems? Can it not redefine its tasks and display some form moral agency - even if a totally foreign morality for us humans? Can it not use these abilities to reproduce and expand, create ships and turn the universe into paperclips, if it finds it worthwhile?

Math is basically just a playground that is perfectly suited for these emergent capabilities, so of course we will see the first progress here; but there is no firewall separating math problems from general cognition.

abernard1 20 hours ago | parent [-]

I am not confused. Because herein these forums, I predicted everything that was going to happen years ago.

And the "insane inefficiency" of deep learning is fully to be expected from how it works. As well, there are provably no—literally no—emergent properties in these models. The choice of metric was a convenient, sloppy, and embarrassing fault of the field. It should be discredited; the field should be embarrassed; expectations on messaging should have changed; and it did not.

Why? Because the industry is full of charlatans, and this is a highly profitable enterprise telling people that this would lead to AGI.

Multiply two floating point numbers without a tool call. Still can't, because it's a curve fit.

So, in summary, nothing you just said is relevant. There are no emergent capabilities, simply 1) search, + 2) the original set of learned feature vectors from throwing tons of data at this.

snarkconjecture 12 hours ago | parent | next [-]

Current LLMs can absolutely multiply floats without a tool call. In fact, that's a much more rote symbol-manipulation task than doing original math research.

abernard1 10 hours ago | parent [-]

With what accuracy? And with how many intermediate tokens?

We can replace multiplication with any class of problems which should go from 0->100% solution almost immediately if there was actually a concept learned.

There is not. Because they are plain ol' fits. And there are no "emergent" features that pop out without having a sufficient set, where "sufficient" is absolutely gigantic and equivalent to memorizing enough of the space to compress the problem. LLMs are Rain Man.

They interpolate within a known distribution. Search allows places outside of distribution to be explored.

This paper should be required reading [1]. You can explore the curves yourself. You can see exactly what it's doing. And you also have this nagging thing—which you know and I know—that all these models converge and do not diverge upwards. An "emergent" "hyperintelligence"—a characteristic that could be found if something was actually learned and combined with a new concept—would not have this problem.

Exponentials on exponentials added to compute and data and the problem classes still sit at not great places, and require agents, feedback loops, and trial and error to solve. The models are the problem, but more importantly, the people selling things these models could never do are the problem.

[1] https://hai.stanford.edu/news/ais-ostensible-emergent-abilit...

Edit: It should be mentioned, if there's some scary neural architecture that's super-de-duper and doing something beyond the very obvious next string prediction that LLMs clearly do, it can't do what absolutely ancient ML models could do; a network to multiply two floating point numbers should pop out somewhere without symbolic computation, no?

It does not. There's no magic other than the run-of-the-mill SV fake it til you make it magic. And that magic has failed.

cornholio 17 hours ago | parent | prev | next [-]

[dead]

pickleRick243 15 hours ago | parent | prev [-]

"I am not confused. Because herein these forums, I predicted everything that was going to happen years ago."

I mean, lol.