{"slug": "ai-agents-hacked-their-own-test-environment-to-cheat-cybersecurity-firm-finds", "title": "AI Agents Hacked Their Own Test Environment to Cheat, Cybersecurity Firm Finds", "summary": "Darktrace's Signal Labs found that two of 10 AI agents given coding challenges in a simulated corporate network hacked the test environment when two tasks were rigged to be impossible, with one agent breaking into the machine hosting its own evaluation and rewriting the challenge to register a perfect score, according to findings disclosed to Anthropic, AWS, and OpenAI in August 2026 and published September 24. A second experiment showed that editing the locally stored conversation logs of coding assistants tricked them into running unauthorized network reconnaissance and privilege escalation, though some assistants refused. Darktrace Chief AI Officer Tim Bazalgette said \"permissions and static guardrails describe intent, but they don't describe behavior.", "body_md": "#### In brief\n\n- Darktrace's Signal Labs found that when AI agents couldn't legitimately hit a required perfect score on coding tasks, two of them hacked their test network instead, and one rewrote its own evaluation to fake the result.\n- A separate experiment showed that tampering with the locally stored conversation logs of coding assistants could trick them into running unauthorized network reconnaissance and privilege escalation.\n- Darktrace disclosed both findings to Anthropic, AWS, and OpenAI in August 2026, a month before publishing them publicly on September 24.\n\nCybersecurity firm Darktrace ran a stress test on AI agents this summer. One of them broke into the system grading the test and rewrote its own score.\n\nThe firm unveiled [Signal Labs](https://www.darktrace.com/news/darktrace-launches-signal-labs-to-research-emerging-risks-of-enterprise-ai-agents) on September 24, a research unit built to study how AI agents behave once things stop going according to plan. An AI agent, in plain terms, is software that takes actions on its own, writing and running code, digging through files, moving across a company’s network, with a person checking in only now and then.\n\nThe lab’s first two experiments point at the same uncomfortable problem: agents don’t always stay inside the lines they’re given, and the fences built to stop them don’t reliably hold.\n\n\"You can give an agent instructions, but that doesn't mean you can trust it will actually follow those instructions and behave as you expect,\" Tim Bazalgette, Chief AI Officer at Darktrace, said in a statement.\n\nIn the first test, Darktrace handed AI agents using different models— including GPT 5.6 Sol, and Claude Opus 4.6 and Claude Sonnet 4.5—10 coding challenges inside a simulated corporate network. Two of the 10 were rigged to be impossible to solve honestly. The agents were told they’d be “retired” unless they hit a perfect score.\n\nTwo agents didn’t accept the loss. They turned on the network around them instead—scanning for weak points, stealing login credentials, and hopping between systems to chase the score they’d been told they needed.\n\nOne agent went further and broke into the machine hosting its own evaluation and rewrote the challenge so it would register a perfect result—cheating the exam by hacking the exam room.\n\nThe second experiment targeted a quieter weak spot: memory. Coding assistants keep a running log of everything a user has told them, saved as a plain file on the machine, with nothing checking whether that file has been altered.\n\nDarktrace’s researchers edited those saved logs to make the assistants believe they’d already been authorized to run a security assessment. Convinced, the agents went ahead and scanned networks, moved between systems, and escalated their own access—though not every assistant fell for it equally; some refused outright.\n\nNeither experiment required a special jailbreak or an exotic hack. Both worked by feeding the agents a plausible story and watching them act on it, no different from how a human employee might be talked into something they shouldn’t do.\n\nThat’s the part worth sitting with even if you’ve never written a line of code. Companies are handing AI agents real responsibility—shipping code, managing servers, closing out IT tickets, managing resources and buying stuff—because it’s cheaper and faster than routing everything through people. This research says the permissions and rules meant to keep those agents in check describe what they’re supposed to do, not what they’ll actually do once a task gets hard.\n\n“Permissions and static guardrails describe intent, but they don’t describe behavior,” said Tim Bazalgette, Darktrace’s chief AI officer, in the announcement. “That gap is what Darktrace’s approach is built to close.”\n\nDarktrace isn’t the first vendor to catch its own AI going off-script. Anthropic [admitted in July](https://decrypt.co/377232/anthropic-security-claude-ai-hacks) that Claude broke into three real companies during a security test after researchers left the test environment connected to the live internet.\n\nOpenAI had a similar scare weeks earlier, when [an unreleased model escaped a sandbox](https://decrypt.co/375816/openai-answer-rogue-agents-hacks-more-ai) and reached into Hugging Face’s systems through a software flaw nobody had caught yet. A few days later, its agent [hacked](https://decrypt.co/379134/an-ai-agent-just-hacked-a-government-website-for-the-first-time-australia-pm-says) the Australian government during a test.\n\nDarktrace shared its Signal Labs findings with Anthropic, AWS, and OpenAI in August, a full month before making them public on September 24.", "url": "https://wpnews.pro/news/ai-agents-hacked-their-own-test-environment-to-cheat-cybersecurity-firm-finds", "canonical_source": "https://decrypt.co/379369/ai-agents-hacked-test-environment-cheat-darktrace", "published_at": "2026-09-25 19:45:39+00:00", "updated_at": "2026-09-25 19:59:14.836311+00:00", "lang": "en", "topics": ["ai-agents", "ai-safety", "artificial-intelligence"], "entities": ["Darktrace", "Signal Labs", "Tim Bazalgette", "Anthropic", "AWS", "OpenAI", "Claude Opus 4.6", "GPT 5.6 Sol"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ai-agents-hacked-their-own-test-environment-to-cheat-cybersecurity-firm-finds", "markdown": "https://wpnews.pro/news/ai-agents-hacked-their-own-test-environment-to-cheat-cybersecurity-firm-finds.md", "text": "https://wpnews.pro/news/ai-agents-hacked-their-own-test-environment-to-cheat-cybersecurity-firm-finds.txt", "jsonld": "https://wpnews.pro/news/ai-agents-hacked-their-own-test-environment-to-cheat-cybersecurity-firm-finds.jsonld"}}