In plain words: Turing's original three-player guessing game was run as he described it: a judge chats to tell machine from human, with a man-imitates-woman round as a check. Only one judge was fooled by the language model, so claims that machines pass the test look premature.
Abstract · A Rigorous Turing Test: a Foundation for Evaluating Artificial General Intelligence
Several studies claim that large language models have passed the Turing Test and hence can "think", yet none follow Turing's original instructions precisely. Passing the test holds significance as evidence that a machine demonstrates human-like intelligence, and as a marker for artificial-general intelligence in commercial and legal domains. We conducted Turing's three-player imitation game with an LLM by following the guidelines identified by Turing and applying scientific standards wherever detailed instructions were missing. We performed a computer-imitates-human game without duration constraints and a man-imitates-woman game as a benchmark. In the computer-imitates-human game, only one participant misidentified the large language model, indicating that claims of large language models' passing the Turing test are premature. Participants required over five minutes for both tasks, with the man-imitates-woman game taking longer; shorter time limits elsewhere may explain earlier positive results. We expect the Turing Test to remain central in assessing machine intelligence for years to come.
Sharon Temtsin, Diane Proudfoot, David Kaber, Christoph Bartneck
arXiv:2501.17629 · cs.HC, cs.AI, cs.CY · submitted Jan 29, 2025 · updated Aug 8, 2026
abstract · pdf · html