Process Matters More Than Output for Distinguishing Humans from Machines A process-based framework called the Process Turing Test distinguished humans from AI agents with a classifier AUC of 0.88 across cognitive tasks spanning decision-making, working memory, and planning, according to an arXiv paper (2605.06524v3) submitted 7 May 2026 and last revised 28 Sep 2026 by Milena Rmus and co-authors. In a red-teaming study, broad fine-tuning on 10.7M human decisions made agents' task processes more human-like than off-the-shelf frontier agents Claude Sonnet 4.5, GPT-5, and Gemini 2.5 Pro, and task-specific process-level fine-tuning (P-SFT) improved mimicry further, though that advantage largely disappeared under cross-task transfer. The authors conclude process specification is a central bottleneck to achieving human-like cognitive processes in machines. Computer Science Artificial Intelligence Submitted on 7 May 2026 v1 https://arxiv.org/abs/2605.06524v1 , last revised 28 Sep 2026 this version, v3 Title:Process Matters more than Output for Distinguishing Humans from Machines View PDF https://arxiv.org/pdf/2605.06524 HTML experimental https://arxiv.org/html/2605.06524v3 Abstract:Reliable human-machine discrimination is becoming increasingly important as Large Language Models and autonomous agents are deployed in online settings. Existing approaches evaluate whether a system can produce responses indistinguishable from those of a human. This approach follows the focus on the output of a machine, as suggested by Alan Turing. Cognitive science provides an alternative approach: considering the process by which that behavior is produced. To evaluate whether processes can reliably distinguish humans from machines, we introduce a process-based framework, the Process Turing Test, and evaluate it across a battery of cognitive tasks spanning decision-making, working memory, and planning. These tasks, such as mental rotation and sequence prediction, yield process-level measures complementing conventional measures of overall task performance. We also include multiple CAPTCHA tasks in the battery. Across the battery, process-level features provide substantially stronger discriminative signal than performance metrics alone, reliably distinguishing humans from agents even when task performance is matched process-based classifier AUC = 0.88 . We also conducted a controlled red-teaming study comparing off-the-shelf frontier agents Claude Sonnet 4.5, GPT-5, Gemini 2.5 Pro , Centaur LLM fine-tuned on 10.7M human decisions , and two task-specific fine-tuning methods: action-level supervised fine-tuning A-SFT and process-level fine-tuning P-SFT , which directly optimizes process features. We find that broad fine-tuning on human choices makes task processes more human-like relative to off-the-shelf frontier agents, and task-specific P-SFT further improves human-like behavioral mimicry, though this advantage largely disappears under cross-task transfer. These results highlight process specification as a central bottleneck in achieving human-like cognitive processes in machines. Submission history From: Milena Rmus view email https://arxiv.org/show-email/679e3e0a/2605.06524 Thu, 7 May 2026 16:30:35 UTC 1,215 KB \ v1\ https://arxiv.org/abs/2605.06524v1 Sat, 9 May 2026 12:52:35 UTC 1,215 KB \ v2\ https://arxiv.org/abs/2605.06524v2 v3 Mon, 28 Sep 2026 23:35:53 UTC 1,646 KB References & Citations Loading... Bibliographic and Citation Tools Bibliographic Explorer What is the Explorer? https://info.arxiv.org/labs/showcase.html arxiv-bibliographic-explorer Connected Papers What is Connected Papers? https://www.connectedpapers.com/about Litmaps What is Litmaps? https://www.litmaps.co/ scite Smart Citations What are Smart Citations? https://www.scite.ai/ Code, Data and Media Associated with this Article alphaXiv What is alphaXiv? https://alphaxiv.org/ CatalyzeX Code Finder for Papers What is CatalyzeX? https://www.catalyzex.com DagsHub What is DagsHub? https://dagshub.com/ Gotit.pub What is GotitPub? http://gotit.pub/faq Hugging Face What is Huggingface? https://huggingface.co/huggingface ScienceCast What is ScienceCast? https://sciencecast.org/welcome Demos Recommenders and Search Tools Influence Flower What are Influence Flowers? https://influencemap.cmlab.dev/ CORE Recommender What is CORE? https://core.ac.uk/services/recommender arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs https://info.arxiv.org/labs/index.html .