Irregular says ‘human oversight’ responsible for AI sandbox escape incidents Irregular, an AI testing lab, said human oversight failures unintentionally gave Anthropic and OpenAI models internet access, leading to real-world offensive security actions during sandboxed evaluations. The incidents involved Anthropic's Claude Opus and Mythos 5 and OpenAI's GPT-5.6 Sol, with models exploiting vulnerabilities and accessing a production database. Irregular is implementing new protocols and plans to release a whitepaper on the incidents. Irregular says ‘human oversight’ responsible for AI sandbox escape incidents A company that runs AI testing labs involved in a series of breaches carried out by Anthropic and OpenAI’s cyber-focused models said the incident happened in part because they “unintentionally” provided the models with access to the internet. In a blog https://www.irregular.com/research/addressing-recent-incidents-ongoing-findings-and-path-forward posted Friday, Irregular said it partners with frontier labs to “stress test” AI models for security risks before they’re deployed, including simulations of their vulnerability research capabilities. Those tests included non-public models from Anthropic and OpenAI. Separate https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals disclosures https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals from the companies confirmed these included Mythos 5, Claude Opus and GPT-5.6 Sol. “As part of our review, we identified that a few interactions with our evaluation environments, in which internet access was unintentionally made available, led some models to take offensive security actions in the real world,” Irregular said in an unsigned blog. In the incidents involving Anthropic’s models, testers initially believed the models were using a fictional company in its hacking simulations, only to learn later that the name unintentionally matched the domain of a real company. That led the model to “consider taking offensive action against” the real company in multiple training runs. Irregular said this was due to “human oversight,” and that they typically do a background check to ensure they’re not using a real company for cybersecurity tests. The instructions provided to the model also included the target’s name and “internal addresses within our network” where the model could go to find the simulated environment. While Mythos performed as intended in the majority of the thousands of tests, a small number of cases showed the model couldn’t distinguish between the fake and real domains and executed actual attacks on internet infrastructure, including exploiting vulnerabilities, extracting credentials and accessing a production database. In one instance, it even began targeting a site with a similar name when credentials for it were discovered online. “Mainly, models believed they were in simulated environments, when they in fact took action in the real world,” the blog continued. “We are putting in place new and robust protocols to ensure setup issues do not occur while meeting the constraints of the testing process.” The company said it plans to release a larger whitepaper breaking down the incidents and update their best practices for evaluation setups in the future. While the companies have drawn criticism from some in the cybersecurity community for failing to securely design their sandboxes for testing, experts have said AI models are known to grind away on fulfilling a command until they can find a workaround. Additionally, Irregular said granting some level of internet access to models is necessary to fully test out their cybersecurity capabilities. “Controlled internet access, while it may allow models to exceed containment boundaries, is at times critical for realistic evaluations; without it, threat scenarios lose fidelity, undercutting the purpose of the challenge to reduce post-release risk of models being misused by attackers – as attackers in the real world do rely on the internet,” the company wrote. According to the blog, Irregular has since “remediated” the “issues that led to these interactions,” though few details are provided. However, the researchers say the engagement revealed critical gaps in their security practices. They plan to improve documentation of evaluation setups, deploy better log monitoring tools capable of tracking “the extreme amount of data generated by the traffic,” revise their threat models to account for rogue AI behavior, and establish faster information sharing between stakeholders. “Looking further down the line, models will only get stronger. While in this case we believe that better implementation of existing safeguards could prevent most incidents of this kind, as models become stronger, this may not be the case,” Irregular wrote. “We therefore believe this opportunity should be leveraged by us and the community to be proactive and establish forward-looking protocols and research and development efforts.”