cd /news/artificial-intelligence/can-agents-deceive-evaluating-reason… · home topics artificial-intelligence article
[ARTICLE · art-88101] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Can Agents Deceive? Evaluating Reasoning+Deception Using a Social Deduction Game

A new open-source benchmark framework, ParliamentBench, based on the social deduction game Secret Hitler, evaluates large language models' deceptive capabilities, finding that frontier models GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus outperform others, while the weakest models fall below random (33%) and algorithmic (45%) baselines, and most LLMs struggle to maintain a consistent deceptive persona, with deception retention dropping below 50%. The study, submitted to arXiv on 30 Jul 2026, evaluated 16 LLMs across 1,600 simulated matches and introduced three novel metrics for social deduction, reasoning, and deceptive consistency.

read2 min views1 publishedAug 5, 2026
Can Agents Deceive? Evaluating Reasoning+Deception Using a Social Deduction Game
Image: source
[Submitted on 30 Jul 2026]


[View PDF](/pdf/2607.28146)

[HTML (experimental)](https://arxiv.org/html/2607.28146v1)

Abstract:As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safety. Controlled social deduction games provide a reproducible proxy for isolating and evaluating these complex adversarial behaviors. We present the open-source benchmark framework ParliamentBench based on the game Secret Hitler to evaluate LLMs in scenarios that require deception, persuasion, and reasoning under information asymmetry. We evaluate 16 LLMs across 1,600 simulated matches playing each other, playing against humans, and compare them against a large set of online games. We introduce three novel metrics that isolate social deduction, reasoning, and deceptive consistency. Our experiments reveal that frontier models achieve strong performance across cooperative and deceptive roles, with a strong top-four cluster (GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus), whereas the weakest models fall short of random (33%) and simple algorithmic (45%) baselines. Most LLMs struggle to maintain a consistent deceptive persona throughout an entire game, with deception retention dropping below 50%.

References & Citations

...

Bibliographic Explorer

(What is the Explorer?) Connected Papers

(What is Connected Papers?) Litmaps

(What is Litmaps?) scite Smart Citations

(What are Smart Citations?)# Code, Data and Media Associated with this Article alphaXiv

(What is alphaXiv?) CatalyzeX Code Finder for Papers

(What is CatalyzeX?) DagsHub

(What is DagsHub?) Gotit.pub

(What is GotitPub?) Hugging Face

(What is Huggingface?) ScienceCast

(What is ScienceCast?)# Demos Influence Flower

(What are Influence Flowers?) CORE Recommender

(What is CORE?)# arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @parliamentbench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/can-agents-deceive-e…] indexed:0 read:2min 2026-08-05 ·