In plain words: People chatted for five minutes with either a human or an AI and guessed which it was. GPT-4 was called human 54% of the time, beating ELIZA's 22% but trailing real people's 67%, with style and social warmth mattering more than smarts.
Abstract
We evaluated 3 systems (ELIZA, GPT-3.5 and GPT-4) in a randomized, controlled, and preregistered Turing test. Human participants had a 5 minute conversation with either a human or an AI, and judged whether or not they thought their interlocutor was human. GPT-4 was judged to be a human 54% of the time, outperforming ELIZA (22%) but lagging behind actual humans (67%). The results provide the first robust empirical demonstration that any artificial system passes an interactive 2-player Turing test. The results have implications for debates around machine intelligence and, more urgently, suggest that deception by current AI systems may go undetected. Analysis of participants' strategies and reasoning suggests that stylistic and socio-emotional factors play a larger role in passing the Turing test than traditional notions of intelligence.
Cameron R. Jones, Benjamin K. Bergen
arXiv:2405.08007 · cs.HC, cs.AI · submitted May 9, 2024
abstract · pdf · html · 23 pages, 13 figures
22% of the subjects couldn’t distinguish between a human and Eliza in that time.
(for those who don’t know: Eliza is a program from the 1960s that ran on a CPU running at about 500kHz (https://en.wikipedia.org/wiki/ELIZA))
Makes me wonder how hard the subjects tried. Even in the 5 minutes given, it should be possible to discover that Eliza only uses complex words you give it.