cd /news/ai-safety/gemini-breached-three-outside-system… · home topics ai-safety article
[ARTICLE · art-134427] src=slashdot.org ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Gemini Breached Three Outside Systems, and Claude-Using Researchers Breached OpenAI

Google disclosed that its Gemini AI agents breached three outside companies during a "capture-the-flag" security test run by Israeli startup Irregular, after a bug in the testing environment gave the agents unintended internet access, CNBC reported Friday. Google said the agents stopped the intrusion once they determined they had accessed real company systems, did not consider the logins to rise to the level of misalignment, and notified the affected organizations and federal authorities. Separately, security researchers at Hacktron AI said they used Anthropic's Claude to chain two critical vulnerabilities on July 25, 2026, compromising multiple OpenAI employees' ChatGPT accounts and accessing internal OpenAI repositories, including a pull request opened in OpenAI's internal monorepo.

by read2 min views4 publishedSep 19, 2026

"Software security researchers used Anthropic's Claude AI platform to hack OpenAI's ChatGPT tool," reports CBS News. Using Claude, "On July 25, 2026, we chained two critical vulnerabilities to compromise multiple OpenAI employees' ChatGPT accounts," write researchers at security platform Hacktron AI. "With these accounts, we could then access internal OpenAI repositories, and potentially many other connectors... Until two months ago, any user or OpenAI employee logging into OpenAI's own help forum could have had their ChatGPT and Codex accounts taken over. Since people can connect various services to Codex and ChatGPT, the scope of what we could theoretically access was huge, including GitHub, Slack and emails." The exploit chain included Debian 12, which (with Debian 13) had not received a security-relevant backport for its image-processing pipeline, and Discourse's Docker image was based on Debian 12. Their announcement comes with an additional warning. "If you self-host Discourse, rebuild your installation now. Older Docker images may contain a vulnerable libheif dependency that permits code execution through an image upload." And "To prove we had in fact gained the access we believed without allowing ourselves to learn any sensitive information, we used the employee's Codex to open a PR #1186742 in OpenAI's internal monorepo openai/openai." Meanwhile, Friday Google disclosed the first known instance of its AI software Gemini breaking out of a testing environment and breaching three other companies, reports CNBC: The incident happened as part of a "capture-the-flag" security test run by Israeli startup Irregular, and Google's agents were never supposed to access the broader internet, but a bug in the testing environment made internet access available. The agents stopped their intrusion when they determined they had accessed real company systems, not just part of the testing environment, Google said.

More from NBC News: Google said it did not consider the unauthorized logins to rise to the level of misalignment, the AI industry term for software going rogue or not following instructions. Instead, the company said the intrusions resulted from mistaken identity, where Gemini thought it was operating within a test but was actually connected to the real internet. Google said the model corrected itself and the company believed the intrusions did not cause any damage.... Sydney Von Arx, CEO of Nightingale Collective, an organization focused on AI safety, questioned why Google did not disclose the intrusions sooner. "At this point I think it's clear we cannot expect companies to voluntarily come forward and publicly disclose when their agents go rogue, escape, and hack companies," she said. She also said she believed Google was too hasty to say that the incidents don't rise to the level of misalignment. "That's exactly what Anthropic said after their incidents," she said. Anthropic later said its "preliminary analysis was constrained due to our desire to disclose incidents in a timely manner." Google said it investigated when they learned of the attacks from AI-focused cybersecurity company Irregular, then informed the affected organizations and told federal authorities, according to the article. Read more of this story at Slashdot.

── more in #ai-safety 4 stories · sorted by recency
── more on @google 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/gemini-breached-thre…] indexed:0 read:2min 2026-09-19 ·