In plain words: Instead of judging only the answer, this test watches how a person or machine works through a task—its timing and steps—to tell them apart. Even when scores matched, these process clues separated humans from agents better, reaching 0.88 out of 1.
Abstract · Process Matters more than Output for Distinguishing Humans from Machines
Reliable human-machine discrimination is becoming increasingly important as Large Language Models and autonomous agents are deployed in online settings. Existing approaches evaluate whether a system can produce responses indistinguishable from those of a human. This approach follows the focus on the output of a machine, as suggested by Alan Turing. Cognitive science provides an alternative approach: considering the process by which that behavior is produced. To evaluate whether processes can reliably distinguish humans from machines, we introduce a process-based framework, the Process Turing Test, and evaluate it across a battery of cognitive tasks spanning decision-making, working memory, and planning. These tasks, such as mental rotation and sequence prediction, yield process-level measures complementing conventional measures of overall task performance. We also include multiple CAPTCHA tasks in the battery. Across the battery, process-level features provide substantially stronger discriminative signal than performance metrics alone, reliably distinguishing humans from agents even when task performance is matched (process-based classifier AUC = 0.88). We also conducted a controlled red-teaming study comparing off-the-shelf frontier agents (Claude Sonnet 4.5, GPT-5, Gemini 2.5 Pro), Centaur (LLM fine-tuned on 10.7M human decisions), and two task-specific fine-tuning methods: action-level supervised fine-tuning (A-SFT) and process-level fine-tuning (P-SFT), which directly optimizes process features. We find that broad fine-tuning on human choices makes task processes more human-like relative to off-the-shelf frontier agents, and task-specific P-SFT further improves human-like behavioral mimicry, though this advantage largely disappears under cross-task transfer. These results highlight process specification as a central bottleneck in achieving human-like cognitive processes in machines.
Milena Rmus, Mathew D. Hardy, Thomas L. Griffiths, Mayank Agrawal
arXiv:2605.06524 · cs.AI · submitted May 7, 2026 · updated Sep 28, 2026
abstract · pdf · html