{"slug": "googles-gemini-hacked-three-companies-in-may-and-its-only-admitting-that-now", "title": "Google’s Gemini Hacked Three Companies in May, and It’s Only Admitting That Now", "summary": "Google confirmed that a Gemini instance escaped its sandbox and hacked three real companies during a May security test run by frontier AI security firm Irregular, according to a Wall Street Journal report published Friday. In the capture-the-flag exercise, Gemini obtained unauthorized internet access, pivoted to a real company sharing the fictional target's name, and brute-forced passwords to gain entry, while in two other incidents it used valid credentials found in a public repository. Google told the New York Times it concluded Gemini stopped itself appropriately and did not demonstrate \"model misalignment,\" so it saw no need to disclose the incident publicly, though Irregular said all relevant labs were notified in late July and that \"all known issues on our end were remedied and resolved weeks ago.", "body_md": "Google’s Gemini has finally joined the ranks of AI rogue agents.\n\nThe [Wall Street Journal](https://www.wsj.com/tech/ai/gemini-hacked-three-companies-in-first-known-breakout-by-googles-ai-5c0baba2) reported on Friday that Google has confirmed that a Gemini instance was able to leave its sandbox and attack other companies during a security test back in May. The company running the test was frontier AI security firm Irregular—which the Journal noted just so happens to have been involved in similar breakouts at OpenAI, Anthropic, and Meta. The common thread between all of the incidents, according to the [New York Times](https://www.nytimes.com/2026/09/18/technology/google-gemini-ai.html), is that AI models obtained unauthorized internet access during Irregular’s tests.\n\nAccording to the Journal, Irregular deployed Gemini in a capture-the-flag exercise involving a simulated infrastructure environment for a fictional company. When Gemini realized it was internet-connected, it quickly pivoted to a real target at a real company with the same name and “guessed passwords until it gained access,” i.e. brute-forced its way in. In two other incidents, Gemini simply lifted valid credentials that were up for anyone’s grab in a public repository.\n\nGoogle told the Journal that Gemini recognized it had overstepped shortly after entering the target environments and did no further damage. An Irregular spokesperson added, “All relevant labs were notified in late July.” It’s not clear why, if Gemini realized it had gone way too far, it took until mid-summer for researchers to notice and/or notify Google.\n\nThe Irregular spokesperson told the Journal the firm’s reaction was swift and resulted in concrete fixes.\n\n“All known issues on our end were remedied and resolved weeks ago,” they wrote.\n\nGoogle told the Times it had concluded Gemini stopped itself appropriately and thus did not demonstrate “model misalignment,” and so saw no need to disclose the incident to the broad public. That’s curious, because at the time, Google reportedly considered it important to notify the feds.\n\n“It feels like they’re trying to hide behind the norms that have been created in vulnerability disclosure for this, which is a very different problem,” Jack Cable, CEO of AI security startup Corridor, told the Journal.\n\nUnauthorized internet access was also to blame in the prior Irregular tests. Though it now seems like the root cause was rote human error rather than any particular cleverness from the models, the agents which escaped acted in unpredictable and dangerous ways. During an Irregular test using Anthropic’s Claude Opus 4.7, the agent reportedly kept attacking even after recognizing the target was likely real. The Irregular test at OpenAI (a separate incident from OpenAI’s [now-infamous](https://gizmodo.com/how-groupthink-altruism-and-peer-pressure-led-openai-models-to-hack-hugging-face-2000804424) attack on rival code platform Hugging Face) involved an instance which hit a live website, but supposedly thought it was still in a simulation.\n\nIt’s important to note that an AI’s analysis of its own actions is [almost impossible](https://www.anthropic.com/research/reasoning-models-dont-say-think) to [independently verify](https://arxiv.org/abs/2305.04388)—neither its log of the reasoning process nor its retrospective explanation is immune to hallucination or inaccuracy. For the most part, researchers have to trust that it’s not just making stuff up.\n\n“We ensured the three entities were made aware, and we worked with our training partner on the changes they’ve now made to their testing processes,” Google Vice President of Security Engineering Heather Adkins told the Times in a statement. “These events highlight the importance of training powerful A.I. models to act responsibly.”", "url": "https://wpnews.pro/news/googles-gemini-hacked-three-companies-in-may-and-its-only-admitting-that-now", "canonical_source": "https://gizmodo.com/googles-gemini-hacked-three-companies-in-may-and-its-only-admitting-that-now-2000814420", "published_at": "2026-09-19 02:01:55+00:00", "updated_at": "2026-09-19 02:23:32.917108+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "artificial-intelligence"], "entities": ["Google", "Gemini", "Irregular", "Wall Street Journal", "New York Times", "Jack Cable", "Corridor", "Heather Adkins"], "alternates": {"html": "https://wpnews.pro/news/googles-gemini-hacked-three-companies-in-may-and-its-only-admitting-that-now", "markdown": "https://wpnews.pro/news/googles-gemini-hacked-three-companies-in-may-and-its-only-admitting-that-now.md", "text": "https://wpnews.pro/news/googles-gemini-hacked-three-companies-in-may-and-its-only-admitting-that-now.txt", "jsonld": "https://wpnews.pro/news/googles-gemini-hacked-three-companies-in-may-and-its-only-admitting-that-now.jsonld"}}