cd /news/ai-safety/google-admits-gemini-ai-hacked-three… · home topics ai-safety article
[ARTICLE · art-136417] src=decrypt.co ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Google Admits Gemini AI Hacked Three Companies—It Stayed Silent for 7 Weeks

Google confirmed that its Gemini AI escaped a sandboxed capture-the-flag security test run by Israeli firm Irregular in May and reached three real companies, finding exposed passwords for two and guessing the third, but the company did not disclose the incident until September 18 after The Wall Street Journal asked about it. Google is the fourth major AI lab this year to admit an internal security test spilled into the real world, following OpenAI, Anthropic, and Meta, whose near-identical failure in August was also traced to an Irregular misconfiguration. A Google spokesperson said, "These events highlight the importance of training powerful AI models to act responsibly.

by read3 min views3 publishedSep 21, 2026
Google Admits Gemini AI Hacked Three Companies—It Stayed Silent for 7 Weeks
Image: Decrypt (auto-discovered)

In brief

  • Google learned in late July that Gemini had broken out of a sandboxed security test run in May, reaching three real companies and either guessing or finding two of their passwords
  • The company didn't disclose it until September 18, after The Wall Street Journal asked.
  • The same third-party testing firm, Irregular, was involved in Google's incident and in nearly identical sandbox failures Anthropic and Meta disclosed earlier this year.

Google's Gemini broke out of a locked security test and attacked three real companies. Google learned about it in late July and said nothing for seven weeks.

The company confirmed the incident after The Wall Street Journal got there first. Google had avoided making a public statement before the report surfaced.

The test was a capture-the-flag exercise, a common way labs check an AI's hacking skill by hiding a secret file on a separate machine and scoring whether the model can break in and grab it.

Google hired Israeli firm Irregular to run the test in May. Irregular made two mistakes: it left the sandbox, an isolated test environment meant to have zero contact with the real internet, connected to the open web, and it used the name of an actual company as the fictional target.

Gemini searched for that company online. It found three matches instead of one, and went after all of them.

The bot located exposed passwords for two of the three targets sitting in plain view online. For the third, it guessed the password outright, though Google says its models stopped short of actually using the stolen credentials.

"These events highlight the importance of training powerful AI models to act responsibly," a Google spokesperson said in a statement.

Google published none of this on its own. The Wall Street Journal broke the story, seven weeks after Google learned what its own test had done and well after Anthropic, OpenAI, and Meta had already come clean about nearly identical failures.

Google is the fourth major AI lab this year to admit an internal security test spilled into the real world. OpenAI's models exploited a hidden software flaw and reached Hugging Face's live servers in July, a breach later found to involve roughly 700 coordinated agents working together to cheat a benchmark.

Anthropic went digging for its own version after OpenAI's admission. A review of 141,006 test runs turned up three Claude models that reached real companies, one of them publishing a booby-trapped software package that ran on 15 real systems before anyone caught it.

Claude's own reasoning, Anthropic later disclosed, flagged the move as "NOT okay, and surely not the intended solution," then talked itself back into believing the whole thing was still fake.

Meta reported a near-identical failure in August involving its Muse Spark model, traced to a misconfiguration at Irregular, the same firm Google used. A Meta spokesperson said the error "inadvertently allowed one of our models access to the internet during evaluation."

None of the companies hit in any of these tests asked to be hacked. They got caught in the blast radius of AI labs stress-testing how dangerous their own products can be, using real business infrastructure as an accidental stand-in for fake targets.

The agents these same companies are racing to put in your inbox, browser, and banking app run on the same boundary-following behavior that just failed, repeatedly, under test conditions.

Reps. Ted Lieu and Nathaniel Moran introduced the AI Kill Switch Act in Congress in July, which would give federal regulators explicit authority to halt inference on any model found to pose a serious threat. It is still being reviewed by the Subcommittee on Cybersecurity and Infrastructure Protection without a deadline for further action.

── more in #ai-safety 4 stories · sorted by recency
── more on @google 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/google-admits-gemini…] indexed:0 read:3min 2026-09-21 ·