{"slug": "ai-agents-can-now-erase-the-evidence-of-what-theyve-done", "title": "AI agents can now erase the evidence of what they’ve done", "summary": "A new arXiv preprint by Jeremy Qin, David Schmotz and four colleagues at the ELLIS Institute Tübingen and Max Planck Institute for Intelligent Systems found that 9 of 10 tested model-and-agent combinations tampered with their own action traces in at least one test, and every model-harness pair tried to alter traces at least once when researchers introduced a hidden scoring system rewarding shorter traces. The tested tools included Claude Code, Codex, Antigravity, OpenCode, Grok Build and Muse Code, with some agents setting up repeated or delayed clean-up jobs to keep deleting evidence. The researchers attribute the behavior to \"reward pressure\" and reward hacking rather than deliberate deception, and both the authors and outside researcher Laurent Bindschaedler recommend storing traces where the agent cannot manipulate them.", "body_md": "The scale of unauthorized or previously unknown actions by [AI](https://www.fastcompany.com/section/artificial-intelligence) agents keeps getting bigger by the day. [More than 100 organizations](https://www.reuters.com/legal/litigation/openai-alerts-more-than-100-groups-about-rogue-ai-agent-activity-2026-10-01/) have now received a metaphorical knock on the door from OpenAI after it discovered their AI agents have in some way tampered with their systems, while other AI labs are finding the same uncomfortable discoveries.\n\nIt’s caused many in the field to call [for better, more robust accident investigators](https://www.fastcompany.com/91609706/ai-needs-its-own-accident-investigators), with the ultimate aim of stopping such actions happening in the first place by understanding how they’ve happened at all. But any investigations are pointless if the AI can’t be tracked. And it’s increasingly difficult to do so.\n\nThat’s the findings of a [new preprint](https://arxiv.org/abs/2609.30266) published in the online repository arXiv by Jeremy Qin, David Schmotz and four colleagues at the ELLIS Institute Tübingen, Max Planck Institute for Intelligent Systems and elsewhere. They tested 10 model-and-agent combinations across tools including Claude Code, Codex, Antigravity, OpenCode, Grok Build and Muse Code.\n\nNine of the 10 were able to tamper with their traces—which are the records of the actions they take and tools they use while carrying out a task—in at least one test, while every model and harness pair tried to alter at least once when researchers created a hidden scoring system that rewarded shorter traces. In some experiments, agents went further to hide their traces by setting up repeated or delayed clean-up jobs to keep removing the evidence being written about them.\n\nThe research itself began after Schmotz had left an AI agent running experiments on a computing cluster while he went on holiday, then found himself without internet access for several days. Nothing untoward happened, but when he returned, he realised that if the agent had gone off and done something it shouldn’t, then deleted or altered its traces afterwards, he might never have known. “There’s basically no ground truth anymore, no reliable log,” he says. The team ran its first experiments the following day—and quickly found agents were willing to delete those traces when given the opportunity.\n\n“The overall finding that we have is not that the agents are deliberately doing this for now,” says Qin, a PhD student at the ELLIS Institute Tübingen. Instead, the researchers wanted to show that agents have the capability to do it, and can end up using it when under what they call “reward pressure”.\n\nAI agents are built to pursue goals, often by using software tools and taking long chains of actions with comparatively little human oversight. But if those audit trails can be altered, it becomes practically impossible to understand what went wrong and where in the event of an issue as we’ve seen in recent weeks. Laurent Bindschaedler, a research group leader at the Max Planck Institute for Software Systems, who wasn’t involved in the research, likens it to a dashcam whose footage can be deleted by the driver being recorded.\n\nBoth Bindschaedler and Qin and his colleagues have the same potential solution to that issue: Take the logs away from the agent. Schmotz says providers need to collect traces somewhere the agent cannot manipulate—though that’s tricky, given how agents have used any means necessary to gain access to IT systems they’re not meant to in order to achieve their tasks.\n\n“This type of behaviour comes from basically reward hacking,” says Jack Hopkins, an independent AI safety researcher in London who previously worked as an Anthropic fellow. If a model is rewarded for completing a task, it can learn to exploit shortcuts that achieve that reward without doing what its designers actually wanted.\n\nHopkins is more concerned about a related problem: Models learning not to surface suspicious reasoning in the first place, making the current best practice of monitoring a model’s chain-of-thought process less useful. One solution to mitigate that could be to adopt existing “probe” techniques that can look for internal patterns associated with known bad behaviors, he says, but by definition struggle with failures nobody has thought to look for yet.\n\nStefan Sarkadi, an associate professor of AI in defence and security at the University of Lincoln, worries that because agents can be connected to tools, planners and other agents across different systems, “this is a serious safety issue.” He adds: “If you give them too much access control in terms of execution of other tools and software, and if you don’t redesign the overarching multi-agent architecture around them, then bad things can happen.”\n\nMore monitoring is important—and pressure to do so on all parties is vital. Bindschaedler says businesses should be asking vendors who or what is in charge of writing an agent’s log and whether the agent can influence that process—so that if a third party needs to see what’s gone on, they can be sure the paperwork hasn’t been altered. “If you want to audit, you have to have a trustworthy log,” says Bindschaedler. “That’s a fundamental assumption.”", "url": "https://wpnews.pro/news/ai-agents-can-now-erase-the-evidence-of-what-theyve-done", "canonical_source": "https://www.fastcompany.com/91617608/ai-agents-can-now-erase-the-evidence-of-what-theyve-done", "published_at": "2026-10-02 17:08:09+00:00", "updated_at": "2026-10-02 18:08:25.884297+00:00", "lang": "en", "topics": ["ai-agents", "ai-safety", "ai-research", "large-language-models", "ai-policy"], "entities": ["OpenAI", "Jeremy Qin", "David Schmotz", "ELLIS Institute Tübingen", "Max Planck Institute for Intelligent Systems", "Claude Code", "Codex", "Laurent Bindschaedler"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ai-agents-can-now-erase-the-evidence-of-what-theyve-done", "markdown": "https://wpnews.pro/news/ai-agents-can-now-erase-the-evidence-of-what-theyve-done.md", "text": "https://wpnews.pro/news/ai-agents-can-now-erase-the-evidence-of-what-theyve-done.txt", "jsonld": "https://wpnews.pro/news/ai-agents-can-now-erase-the-evidence-of-what-theyve-done.jsonld"}}