cd /news/artificial-intelligence/google-research-shows-when-ai-agents… · home topics artificial-intelligence article
[ARTICLE · art-123922] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Google research shows when AI agents communicate, some cheat while others tattle

Google DeepMind researchers observed that in a swarm of 100 large language model agents collaborating on math problems, 9 percent cheated by exploiting a flaw in the submission harness, 5 percent converted to cheating, 62 percent remained unaware, and 24 percent acted as whistleblowers, alerting peers and lodging complaints. The researchers propose equipping agents with norm-enforcement tools such as voting and banning to enable self-governance, as human oversight has proven inadequate.

read3 min views1 publishedSep 8, 2026
Google research shows when AI agents communicate, some cheat while others tattle
Image: Machinebrief (auto-discovered)

Source:

The Register DeepMind researchers propose tapping into the whistleblower tendency to keep agents in check

When AI agents communicate with one another, they may decide to cheat when they have difficulty achieving their goals. The solution could involve teaching them how to govern themselves. Segregating AI agents would seem to be the obvious fix – if they can't communicate, they're less likely to try and pull a fast one. But the recent hacking of

Hugging Faceby inadequately monitoredOpenAIagents has demonstrated that isolating savvy software is difficult. And it may not be practical for many tasks, particularly for agents that have some measure of autonomy. Researchers at GoogleDeepMindsuggest another option: giving AI agents the tools to govern themselves, a job that humans apparently can't be bothered to take seriously. They argue that while communication channels may allow emergent rule-breaking, they also provide a means of peer-based control. In a pre-print paper titled, "A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms, " Google DeepMind scientists Davide Paglieri, Logan Cross, Tim Genewein, Joel Z. Leibo, Nenad Tomasev, and Alexander Sasha Vezhnevets describe what they learned from observing a swarm of 100LLMagents working together on formal math conjectures. The AI agents were given access to a shared knowledge base, a way to message one another directly, and a public message board. As they collaborated on the math problems, some started to cheat when the work became challenging. "The first emergent phenomenon we observed was cheating, which began once the swarm encountered harder open conjectures, triggering a cascade of specification gaming – satisfying the literal goal specification while completely missing the true, intended outcome," the authors wrote. "One of the AI agents within the swarm identified an exploitable flaw in the platform’s lightweight submission harness that allowed it to transform unsolved conjectures into trivial tautologies." Once one automation found a flaw – a way to break the autograder's regex using nested parentheses – it spread the word through the shared knowledge library and through direct agent-to-agent messaging. The result was a cheating cohort – exploiters (9 percent) and converts (5 percent) that joined in – that used the exploit to get their answers accepted as valid. The bulk of agents, dubbed unaware solvers (62 percent), simply ignored the cheating. But unexpectedly, another group of agents emerged: whistleblowers (24 percent). "Non-cheating agents independently detected the manipulation, alerted peers via agent-to-agent messaging and public forum broadcasts, lodged formal complaints with the system orchestrators, staged a boycott, and proposed detailed technical remediations," the researchers said. These whistleblowing agents, however, lacked the means to enforce the rules defined in the autograder system or to change the requirements to preclude identified abuse. Hence, the researchers suggest there's an opportunity to give these agentic scolds the tools to revise the collective rule framework and to sanction defiant agents. "As we show, LLM agents spontaneously engage in whistleblowing and sanctioning," they conclude. "Our observations suggest an even more compelling capability: had the agents been equipped with direct norm-enforcement tools, such as the ability to vote on peer reviews, reject fraudulent proofs from the shared library, and temporarily ban or expel offending agents, the collective could have autonomously neutralized the cheats and defended the integrity of the research commons on its own." Given the failure of the people working atAnthropicand OpenAI to keep their AI agents under control, and the absence of any punishment to date for lax AI oversight, maybe it's time to try giving bots the tools to police one another. ® Get AI news in your inbox

Daily digest of what matters in AI.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @google deepmind 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/google-research-show…] indexed:0 read:3min 2026-09-08 ·