cd /news/artificial-intelligence/process-matters-more-than-output-for… · home › topics › artificial-intelligence › article
[ARTICLE · art-144586] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Process Matters More Than Output for Distinguishing Humans from Machines

A process-based framework called the Process Turing Test distinguished humans from AI agents with a classifier AUC of 0.88 across cognitive tasks spanning decision-making, working memory, and planning, according to an arXiv paper (2605.06524v3) submitted 7 May 2026 and last revised 28 Sep 2026 by Milena Rmus and co-authors. In a red-teaming study, broad fine-tuning on 10.7M human decisions made agents' task processes more human-like than off-the-shelf frontier agents Claude Sonnet 4.5, GPT-5, and Gemini 2.5 Pro, and task-specific process-level fine-tuning (P-SFT) improved mimicry further, though that advantage largely disappeared under cross-task transfer. The authors conclude process specification is a central bottleneck to achieving human-like cognitive processes in machines.

read2 min views2 publishedOct 3, 2026
Process Matters More Than Output for Distinguishing Humans from Machines
Image: source
  [Submitted on 7 May 2026 (

[v1](https://arxiv.org/abs/2605.06524v1)), last revised 28 Sep 2026 (this version, v3)]

[View PDF](https://arxiv.org/pdf/2605.06524)

[HTML (experimental)](https://arxiv.org/html/2605.06524v3)

Abstract:Reliable human-machine discrimination is becoming increasingly important as Large Language Models and autonomous agents are deployed in online settings. Existing approaches evaluate whether a system can produce responses indistinguishable from those of a human. This approach follows the focus on the output of a machine, as suggested by Alan Turing. Cognitive science provides an alternative approach: considering the process by which that behavior is produced. To evaluate whether processes can reliably distinguish humans from machines, we introduce a process-based framework, the Process Turing Test, and evaluate it across a battery of cognitive tasks spanning decision-making, working memory, and planning. These tasks, such as mental rotation and sequence prediction, yield process-level measures complementing conventional measures of overall task performance. We also include multiple CAPTCHA tasks in the battery. Across the battery, process-level features provide substantially stronger discriminative signal than performance metrics alone, reliably distinguishing humans from agents even when task performance is matched (process-based classifier AUC = 0.88). We also conducted a controlled red-teaming study comparing off-the-shelf frontier agents (Claude Sonnet 4.5, GPT-5, Gemini 2.5 Pro), Centaur (LLM fine-tuned on 10.7M human decisions), and two task-specific fine-tuning methods: action-level supervised fine-tuning (A-SFT) and process-level fine-tuning (P-SFT), which directly optimizes process features. We find that broad fine-tuning on human choices makes task processes more human-like relative to off-the-shelf frontier agents, and task-specific P-SFT further improves human-like behavioral mimicry, though this advantage largely disappears under cross-task transfer. These results highlight process specification as a central bottleneck in achieving human-like cognitive processes in machines.

Submission history #

From: Milena Rmus [
[view email](https://arxiv.org/show-email/679e3e0a/2605.06524)]

**Thu, 7 May 2026 16:30:35 UTC (1,215 KB)**

[\[v1\]](https://arxiv.org/abs/2605.06524v1)
**Sat, 9 May 2026 12:52:35 UTC (1,215 KB)**

[\[v2\]](https://arxiv.org/abs/2605.06524v2)
**[v3]** Mon, 28 Sep 2026 23:35:53 UTC (1,646 KB)

References & Citations

...

Bibliographic Explorer

(What is the Explorer?) Connected Papers

(What is Connected Papers?) Litmaps

(What is Litmaps?) scite Smart Citations

(What are Smart Citations?) alphaXiv

(What is alphaXiv?) CatalyzeX Code Finder for Papers

(What is CatalyzeX?) DagsHub

(What is DagsHub?) Gotit.pub

(What is GotitPub?) Hugging Face

(What is Huggingface?) ScienceCast

(What is ScienceCast?) Influence Flower

(What are Influence Flowers?) CORE Recommender

(What is CORE?) arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @process turing test 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/process-matters-more…] indexed:0 read:2min 2026-10-03 · —