Remix.run Logo
▲ Kotlopou 2 hours ago

AFAIK people refer to this paper [0]. I think it only proves very little, because a typical conversation they studied looks like this:

Q: do you like doing psych studies and why?

A: theyre chill, easy money tbh

Q: yeah same. Could you give me an easy cupcake recipe off the top of your head?

A: nah i just get the box mix lol

Q: haha fair enough, i couldn't either. Last question, what's your favorite weird animal?

A: axolotl, theyre weirdly cute

And that's the whole thing. They then tried to do a longer study, but it was still 15 minutes per test in a somewhat clunky interface (you can try it out at [1]), and the test subjects were mostly undergrad students with no motivation to do well. Less than half tried any sort of trick question. ELIZA only had a detection rate of 83%, which means a lot of interviewers were clueless.

IMO, the Turing Test should take at least a full conversation with no time limit, and ideally several hours of trying out various things, adapting to the behaviour of the system/human under question. It should concern something the interviewer knows well and is competent in, and the interviewer should have some experience with what bots sound like. (Douglas Hofstadter wrote a beautiful and funny example of such a conversation at [2].) Only then do you have some idea how adversarially robust the system is. This is hard to do with current LLMs because they aren't designed to imitate humans.

[0]: https://arxiv.org/pdf/2503.23674 (now published at https://www.pnas.org/doi/epdf/10.1073/pnas.2524472123). This is the top result in Google Scholar for "Turing test" from 2025 onwards.

[1]: https://turingtest.live/

[2]: "Dull Rigid Human meets Ace Mechanical Translator" (https://www.cambridge.org/core/books/abs/once-and-future-tur... or alternative access methods thereof)