Remix.run Logo
▲ bunderbunder 6 hours ago

I remain unconvinced.

The thing about a generative language model that’s trained from a massive but unknown corpus is, it’s practically (if not theoretically) impossible to evaluate the extent to which data leakage contributes to any particular output.

But I would argue that, as things currently stand, “sophisticated engine for approximately querying a pastiche of the results of human reasoning that comprise its training corpus” remains a more parsimonious explanation than “it’s doing actual reasoning” for how this neural network architecture produces the phenomena we’ve been observing.

▲Muromec 6 hours ago | parent [-]

>sophisticated engine for approximately querying a pastiche of the results of human reasoning that comprise its training corpus

Well if the thing can find and fix bugs in something that is using non-mainstream stuff that is surely not in it's training dataset, that's better than a rubber duck already. Whether it has soul is a different question of course.

▲bunderbunder 6 hours ago | parent | next [-]

“A soul”?

As popular as I know the rhetorical tactic is on both sides of these discussions about LLMs, I’d still thank you not to strawman me.

▲phoghed 6 hours ago | parent | prev [-]

Don’t even bother. These people almost always have some goofy ass, non standard, fluid definition of “thinking” or “reasoning” that cannot ever be met.

▲bunderbunder 5 hours ago | parent [-]

My opinion of their reasoning capability is based in part on (proprietary, non-published, only internally peer reviewed) experiments on GPT-series models’ ability to perform a suite of formal and informal inference and deduction tasks.

Perhaps you could argue that “appropriately applies syllogism to arrive at correct conclusions” is too high a bar to set, but I don’t think it would be fair to call it a “goofy-ass”, “non-standard” or “fluid” element of a reasoning capacity assessment.

▲chpatrick 5 hours ago | parent | next [-]

Yeah I guess the Jacobian conjecture must have been disproved without reasoning.

▲bunderbunder 5 hours ago | parent [-]

It’s hard to say. But supposedly the counter example wasn’t found by an agent running in full auto; it came out of a bunch of back and forth with a human operator. Without, in addition to the aforementioned access to currently non-public information about these models, a detailed transcript of the chat sessions leading up to the discovery, it’s hard to ascribe the reasoning steps involved to any source in particular.

Part of my concern here is that simply pointing out that LLMs appear to be performing tasks that can be done through reasoning, and using that in and of itself as evidence of reasoning, is affirming the consequent.

▲retsibsi 2 hours ago | parent | prev [-]

But you're not just saying they are insufficiently good at reasoning, you're saying they're (probably) not "doing actual reasoning". So we need to know how you are defining "actual reasoning".

I don't think the bar for an actual reasoner can possibly be 'always appropriately applies syllogism to arrive at correct conclusions', because in that case nobody in the world is an actual reasoner. And if your bar were 'sometimes appropriately applies syllogism to arrive at correct conclusions', it's hard to understand why the current generation of AIs doesn't meet it; they are clearly capable of doing so, at least to all outward appearances. (Maybe you think their apparently successful demonstrations of reasoning are illusions, but again, you would need to define what counts as "actual reasoning" vs. a superficially convincing simulation of it.)