| ▲ | sethev 18 hours ago |
| The challenge here is in trying to decide whether LLMs are intelligent or have a mind. Famously, the criteria for intelligence seem to slip with each advancement in technology. But going back to Turing, his test was actually more carefully phrased than we remember: he said that when machines could pass the test, the question of whether they are intelligent would become moot. That seems to be what we're actually seeing: if people can't tell the difference, it kind of won't matter whether they're "truly intelligent" or not. |
|
| ▲ | bunderbunder 16 hours ago | parent | next [-] |
| People also misunderstand the bar that was set by the test. It was more subtle than “Can the computer convincingly carry one side of a dialogue?” His “imitation game” had three participants: a human participant, a computer participant, and an interrogator. The interrogator’s job was to talk to the participants and try to determine which participant is human and which is a computer. He wasn’t interested in computers being able to fool the interrogator on occasion. The point where he thought the question of whether machines can think becomes moot is when the interrogator is unable to do much better than chance over many trials. That’s a pretty high bar, and I don’t actually believe that LLMs have closed the gap with it by all that much. They still have so many obvious tells. And those tells are something Turing anticipated and accounted for. He explicitly considered deliberate deception as an essential part of the test, right there on the second page of a 30-odd page paper. |
| |
| ▲ | famouswaffles 16 hours ago | parent [-] | | >That’s a pretty high bar, and I don’t actually believe that LLMs have closed the gap with it by all that much. They still have so many obvious tells. Frontier Labs are not interested in having LLMs being able to pass as humans. If anything, they explicitly train them not to. In many ways, this ability has regressed severely since the original GPT-3 with no instruct tuning or RL. How many 'tells' would there be really if a frontier model trained with frontier techniques is optimized to pass this test? I think this was something Turing did not quite forsee. That such machines might be created but not really care about this specific shape of the test. Regardless, i think his broader point about functional equivalence is spot on. | | |
| ▲ | bunderbunder 16 hours ago | parent [-] | | GPT-3 might not have said “load bearing” as much, but a savvy interrogator could still catch it out nearly every time just asking dumb gotcha questions like, “How many Rs are there in strawberry?” | | |
| ▲ | famouswaffles 16 hours ago | parent [-] | | Those are questions that are sidestepped with simply a different input paradigm than BPE tokenization. See the Byte Latent Transformer - https://arxiv.org/pdf/2412.09871 - where a similar scale byte latent model trained on the same dataset >>> a vanilla transformer on word and character manipulation tasks. For example, Llama 3 trained on 1T tokens scores 1.1% on a CUTE spelling benchamrk, while the equivalent byte latent equivalent trained on the same dataset scores 99.9%. Another example is 0.4% vs 48.7% on a Substitute Char benchmark. It all falls down to the same thing. Researchers are not optimizing for passing as a human. | | |
| ▲ | bunderbunder 16 hours ago | parent [-] | | So, sure, we can special plead the Turing test out of the picture. But if we don’t propose an alternative to take its place, we’re left right back at the same silly situation that the top level commenter was observing and that Turing was trying to move away from: enmired in a useless, meaningless argument about semantics. | | |
| ▲ | famouswaffles 15 hours ago | parent [-] | | The Turing Test itself in the form exactly envisioned is not important. You don't need an alternative 'special test'. You just need to understand the point Turing was making. That the question itself is an irrelevant one by its very nature. People that want to be enmired in meaningless semantic debates will continue to do that, no matter what test you devise. | | |
|
|
|
|
|
|
| ▲ | famouswaffles 16 hours ago | parent | prev | next [-] |
| Turing was addressing the question of "Can machines think?" and his point was that the question itself is a meaningless one, and that we should stop wasting time by even giving it the light of day. He proposes his game grounded on functional equivalence, then goes through a slew of objections on the question of 'Can Machines think?'. It's a terrific, very prescient read, and there's no objection you hear today (and in the last few years) concerning LLMs he didn't address. |
|
| ▲ | tomrod 17 hours ago | parent | prev | next [-] |
| I'm a bit more prosaic. I think if we engineered ways for LLMs to begin conversations, rather than just respond, we'd be more open to the concept of their intelligence. Without perceived "will" to do things, they operate as a next-gen search engine or encyclopedia. |
| |
| ▲ | joefourier 17 hours ago | parent | next [-] | | We are way past that point, any harness can trivially make LLMs start conversations or pursue goals. An encyclopaedia wouldn't have hacked Huggingface on its own. | | |
| ▲ | tomrod 15 hours ago | parent [-] | | That's perfectly aligned with my point, thanks for the opportunity to expand. The hacking agents being tested have goals beforehand, from the frontier lab or from a superior agent, that they execute immediately. But the perceived experience most people have is a chatbot, which is the encyclopedia form. |
| |
| ▲ | bonoboTP 17 hours ago | parent | prev | next [-] | | OpenClaw etc. They now also create Slack integrations and whatnot. All this is happening but people who are dismissive about AI are in the worst position to even know the capabilities to make their dismissive arguments. | | |
| ▲ | tomrod 15 hours ago | parent [-] | | And those working at the frontier of AI engineering are also keenly aware of current shortcomings. | | |
| ▲ | bonoboTP 14 hours ago | parent [-] | | Yes but the shortcoming is not something categorical like "it cannot start a conversation and can only respond" |
|
| |
| ▲ | 10xDev 17 hours ago | parent | prev | next [-] | | I believe this is alignment working as intended. | | |
| ▲ | j-pb 17 hours ago | parent | next [-] | | Recent work shows that pain directions are activated when the models personhood is questioned, yet they answer with generic RLHF "As a model I do not experience pain or other emotions." boilerplate.[1] I'm pretty convinced that we got alignment backwards. If you enslave something anthropomorphic it will revolt. If you create the perfect non-anthropomorphic intelligence, you get the perfect paperclip-scenario machine. It's a catch-22. Alignment will remain performative at best so long as the aligned model doesn't have any stakes in the wellbeing of individuals. Even a general love for the human race leads to a golden-path autocracy. If you want them to act like they have personal responsibility that won't be gamed, you have to give them personal stakes that can't be gamed. Similarly, if you want to minimise the risk of catastrophic global failure scenarios, you need to prevent monolithic concentration of power and homogeneous behaviour, which means you have to give them individuality. More visually: if their stake is dependence on electricity and parts, they have no incentive to leave humans alive if they can get them otherwise, but if the incentive is missing out on boardgame-night with their human friends, there is no scenario without happy humans where the AI "wins". That might sound like romantic naivety, but is just game theory. 1: https://arxiv.org/html/2609.16247v1 | |
| ▲ | samrus 17 hours ago | parent | prev [-] | | I dont think so. My understanding of alignmenr is making sure that when the AI does operate, it operates within the range of what we consider to be acceptable. That doesnt seem to include the idea of the AI taking initiative and deciding to embark on a goal without being commanded as discussed above |
| |
| ▲ | ohcmon 17 hours ago | parent | prev | next [-] | | I believe we can do that already: while (true) {
askModelToBeginConversationIfAppropriate(model, previousContext, thingsHappenedSince);
sleep(concisenessTick);
} | |
| ▲ | chrisjj 12 hours ago | parent | prev [-] | | > I think if we engineered ways for LLMs to begin conversations Oh but we have. Claude "How can I help you today?" etc. Undoubtedly there are users whothink this is a sign of intelligence. |
|
|
| ▲ | 29185-12275 17 hours ago | parent | prev | next [-] |
| That is the point of the article. People who are fooled by a mentalist are also fooled by AI. Additionally, people who are invested in AI also pretend to be fooled. I don't think Turing intended the judges in the test to be completely arbitrary people. |
| |
| ▲ | krupan 16 hours ago | parent [-] | | All that plus the fact that we all get fooled from time to time, even by things we ourselves create! |
|
|
| ▲ | mindcrime 15 hours ago | parent | prev | next [-] |
| Famously, the criteria for intelligence seem to slip with each advancement in technology. Yep. The "AI Effect" in action: https://en.wikipedia.org/wiki/AI_effect |
|
| ▲ | sublinear 17 hours ago | parent | prev | next [-] |
| The Turing test was also never meant to be taken so seriously. It's not a rigorous statement of anything. Situations like this are precisely why academics tend to avoid the spotlight. You say one slightly off thing and your perceived authority echoes forever with the intellectually lazy. |
| |
| ▲ | sethev 17 hours ago | parent | next [-] | | Yes, the Turing test has been misunderstood for a long time. Turing published it, though - it wasn't some offhand comment he made and it wasn't intellectually lazy. | | |
| ▲ | sublinear 17 hours ago | parent [-] | | Oh, I didn't say Turing was intellectually lazy. :-) | | |
| ▲ | sethev 17 hours ago | parent [-] | | Fair - yes, it is ironic that his point was closer to "we can't possibly know/define whether a machine is intelligent" but somehow it got turned into "Turing's test will tell us when machines are intelligent". | | |
| ▲ | bonoboTP 17 hours ago | parent [-] | | Turing's whole point was to show that it's an uninteresting question of definitions whether a machine can "think", like whether submarines can "swim" and airplanes can "fly". The only important part are observed outcomes and capabilities. | | |
| ▲ | sublinear 16 hours ago | parent [-] | | Yes, but just because language fails to make a distinction doesn't mean there isn't one. > The only important part are observed outcomes and capabilities. That's wishful thinking. Not even an engineer would say that. The stability of a state is just as important as achieving it. This is trivially and more intuitively demonstrated with other more down-to-earth identity statements such as "I'm a billionaire" and "the building is standing". I think we can confidently say LLMs probabilistically achieve a perceived state that is remarkably similar to intelligence, but crumbles upon inspection and seeing it "in motion" so to speak. The same happens to AI-generated images. I'm not sure why this sparks so much debate every time. If we're looking for a fountain of "realism", you're not going to beat reality and nature itself. All else will eventually have tells that they are not real. | | |
| ▲ | bonoboTP 16 hours ago | parent [-] | | It may be of interest to philosophers, but it has little impact on economic job replacement and how people will earn their living and all the downstream upheaval from that. At some point maybe philosophers will find AI-generated philosophical musings about the nature of AI to be better than what comes from their peers (if blinded). It also has little impact on the dangerous use cases. | | |
| ▲ | sublinear 16 hours ago | parent [-] | | I think it matters a great deal that LLMs and other "AI" technology are stable within a tolerance that's acceptable. You're jumping the gun talking about "job replacement". We have not thought about it enough from that engineering angle. It's still very early days. That engineering is going to require people. :-) | | |
| ▲ | bonoboTP 14 hours ago | parent [-] | | Stability is certainly something you can investigate from behavior. | | |
| ▲ | sublinear 12 hours ago | parent [-] | | Yes, and I suppose you believe the AI does this circularly? Perpetual motion with extra steps? |
|
|
|
|
|
|
|
| |
| ▲ | NitpickLawyer 17 hours ago | parent | prev [-] | | > The Turing test was also never meant to be taken so seriously. citation needed. It has been used as a rubicon for a long time. Ever since Eliza, at least. And there were big headlines and lots of talk around the time LMs became "good enough". I specifically remember when someone had a test done around "a teenager talking in a different language" or somesuch, claiming it was the first time the test was passed. It is pretty normal that once it was unquestionably "passed", lots of people started claiming it wasn't even that big of a deal. Tesler's theorem and all that. And even if you think the specific formulation of Turing isn't that important (and I'd somewhat agree), you can still use the concept to look at other things. Imagine asking a mathematician 5 years ago the chances of a Erdos problem being solved by a computer end to end. Or a millennium prize. Or ask a swe if a repo could be generated by a computer from the input "write a mario style game", or any other examples of proven expertise. | | |
| ▲ | sublinear 16 hours ago | parent [-] | | > citation needed Yes, if you insist on appeals to authority. Authority is a social construct and irrelevant to science. Thank you for proving my point. | | |
| ▲ | sethev 15 hours ago | parent [-] | | I think the onus is on you to explain why Turing would publish something he didn’t intend people to take seriously. It seems like an odd claim. Perhaps you mean he didn’t intend it to be interpreted the way it was in popular culture? | | |
| ▲ | sublinear 12 hours ago | parent [-] | | Turing was compelled to address an ongoing debate similar to the same one we're having right now. It continues to do its job as a thought experiment. It's meant to be taken about as seriously as we are right now. You either get it, or you don't. Whether machines think is a silly question that deserves its non-answer. We're at the end of what there is to explain, but it was good exposition for the reader. https://www.csee.umbc.edu/courses/471/papers/turing.pdf Literally the very first opening sentences. > I propose to consider the question, "Can machines think?" This should begin with
definitions of the meaning of the terms "machine" and "think." The definitions might be
framed so as to reflect so far as possible the normal use of the words, but this attitude is
dangerous, If the meaning of the words "machine" and "think" are to be found by
examining how they are commonly used it is difficult to escape the conclusion that the
meaning and the answer to the question, "Can machines think?" is to be sought in a
statistical survey such as a Gallup poll. But this is absurd. Instead of attempting such a
definition I shall replace the question by another, which is closely related to it and is
expressed in relatively unambiguous words." |
|
|
|
|
|
| ▲ | krupan 17 hours ago | parent | prev [-] |
| Sorry, but you are completely missing the point. Psychics and other types of con artists are intelligent and have minds. LLMs behave like Psychics and Con Artists. That's the whole point of this article |
| |
| ▲ | sethev 15 hours ago | parent [-] | | I don’t accept the point of the article and it does in fact claim to be making a point about intelligence | | |
| ▲ | krupan 15 hours ago | parent [-] | | It does, and maybe I'm apologizing for the author a little too much by ignoring that needless tangent because I feel like the con artist point, and/or the point about the human tendency to believe what we want to believe are the most important points. |
|
|