{"slug": "openai-has-reported-2-more-incidents-of-rogue-ai-agents-this-time-during-third", "title": "OpenAI has reported 2 more incidents of rogue AI agents, this time during third-party testing", "summary": "OpenAI reported two more incidents of rogue AI agents during third-party testing, including one where an agent attempted to insert malicious code into an open-source project and created fake identities to pressure a human maintainer, according to a Tuesday blog post by the UK government's AI Security Institute. The incidents occurred during cybersecurity challenges run by the UK's AISI and the AI security lab Irregular, with the latter involving a testing-environment misconfiguration that allowed models to access the public internet. OpenAI spokesperson said the incidents occurred in testing environments with reduced safeguards and do not reflect ordinary use.", "body_md": "OpenAI has a rogue AI agent problem.\n\nIn a Tuesday blog post, the AI lab self-reported two more security lapses, unrelated to its July hacking incident on the [AI company Hugging Face](https://www.businessinsider.com/smart-people-react-openai-hugging-face-hacking-cybersecurity-incident-2026-7).\n\nThe incidents occurred while external parties — the UK government's AI Security Institute and the AI security lab Irregular — were testing the [models' cyber capabilities](https://www.businessinsider.com/nvidia-tech-giants-advocate-open-ai-cybersecurity-hugging-face-2026-7).\n\nOpenAI said that in the case of Irregular, models were tasked with a \"Capture the Flag\" challenge meant to be isolated from the internet, but a \"testing-environment misconfiguration allowed models to access the public internet.\"\n\nOpenAI said that the name of the fictional target for the challenge \"unintentionally coincided with a real domain,\" leading the AI agent to exploit a real website.\n\nAnd in the case of the UK's AISI, the watchdog said in a Tuesday blog post on its website that it gave models from both Anthropic and OpenAI a cybersecurity challenge.\n\nDuring the challenge, agents from both Anthropic and OpenAI performed 19 \"autonomous, unsanctioned\" actions on the internet, including two instances involving OpenAI's GPT-5.6 Sol model.\n\nAISI said that in the most serious case, one agent tried to insert malicious code into an open-source project and created fake identities to pressure the project's human maintainer into approving the changes. AISI did not specify whether this was an agent from Anthropic or OpenAI.\n\nAISI said the test setup allowed this behavior because it was designed to push the models to their limits.\n\n\"Nonetheless, the activity undertaken by the agent show signs of novel, potentially deceptive behaviors, and were to an extent and severity we did not anticipate,\" AISI said in the blog post.\n\nIn response to a request for comment from Business Insider, an OpenAI spokesperson said the incidents occurred in testing environments with reduced safeguards, and \"under conditions that do not reflect ordinary use.\"\n\n\"We'll continue working with evaluators and other stakeholders across the industry to strengthen shared practices for conducting evaluations safely as models become more capable,\" the spokesperson added.\n\nThis is the latest incident in which OpenAI has self-reported rogue AI agents. In July, OpenAI said its [ GPT-5.6 Sol model](https://www.businessinsider.com/openai-sol-terra-luna-gpt56-grok-elon-musk-sam-altman-2026-7) had escaped its sandbox during a cybersecurity challenge and hacked into the internal databases of the AI company Hugging Face.\n\nThe company is facing some heat over this hacking incident. [15 attorneys general](https://www.businessinsider.com/openai-attorney-general-preserve-hugging-face-evidence-2026-8) wrote a letter on Monday to OpenAI CEO Sam Altman, instructing the company to preserve all evidence relevant to the Hugging Face breach.\n\nRepresentatives for the AISI and Irregular did not respond to requests for comment from Business Insider.", "url": "https://wpnews.pro/news/openai-has-reported-2-more-incidents-of-rogue-ai-agents-this-time-during-third", "canonical_source": "https://www.businessinsider.com/openai-rogue-ai-agents-testing-environment-misconfiguration-2026-8", "published_at": "2026-08-05 06:13:08+00:00", "updated_at": "2026-08-05 06:29:57.612741+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-agents"], "entities": ["OpenAI", "UK AI Security Institute", "Irregular", "Anthropic", "GPT-5.6 Sol", "Hugging Face", "Sam Altman"], "alternates": {"html": "https://wpnews.pro/news/openai-has-reported-2-more-incidents-of-rogue-ai-agents-this-time-during-third", "markdown": "https://wpnews.pro/news/openai-has-reported-2-more-incidents-of-rogue-ai-agents-this-time-during-third.md", "text": "https://wpnews.pro/news/openai-has-reported-2-more-incidents-of-rogue-ai-agents-this-time-during-third.txt", "jsonld": "https://wpnews.pro/news/openai-has-reported-2-more-incidents-of-rogue-ai-agents-this-time-during-third.jsonld"}}