cd /news/artificial-intelligence/can-agents-deceive-evaluating-reason… · home topics artificial-intelligence article
[ARTICLE · art-81402] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

A new open-source benchmark framework, ParliamentBench, based on the social deduction game Secret Hitler, evaluates whether large language models (LLMs) can deceive, persuade, and reason under information asymmetry. Testing 16 LLMs across 1,600 simulated matches, the study found that frontier models like GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus perform strongly, while weaker models fall below random (33%) and algorithmic (45%) baselines. Most LLMs struggle to maintain a consistent deceptive persona, with deception retention dropping below 50%.

read1 min views1 publishedJul 31, 2026

arXiv:2607.28146v1 Announce Type: new Abstract: As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safety. Controlled social deduction games provide a reproducible proxy for isolating and evaluating these complex adversarial behaviors. We present the open-source benchmark framework ParliamentBench based on the game Secret Hitler to evaluate LLMs in scenarios that require deception, persuasion, and reasoning under information asymmetry. We evaluate 16 LLMs across 1,600 simulated matches playing each other, playing against humans, and compare them against a large set of online games. We introduce three novel metrics that isolate social deduction, reasoning, and deceptive consistency. Our experiments reveal that frontier models achieve strong performance across cooperative and deceptive roles, with a strong top-four cluster (GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus), whereas the weakest models fall short of random (33%) and simple algorithmic (45%) baselines. Most LLMs struggle to maintain a consistent deceptive persona throughout an entire game, with deception retention dropping below 50%.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @parliamentbench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/can-agents-deceive-e…] indexed:0 read:1min 2026-07-31 ·