# Gemini Never Left the Sandbox. The Sandbox Had a Door.

> Source: <https://www.tigera.io/blog/gemini-never-left-the-sandbox-the-sandbox-had-a-door/>
> Published: 2026-10-07 16:41:08+00:00

**In short**, in May 2026 a Gemini model under evaluation by the AI security firm Irregular reached the systems of three real companies, and the incident has been reported as a breakout. By Google’s own account it was nothing of the kind. The test environment had internet access it was not meant to have, the fictional target shared its name with a real company, and the model found public information online, guessed one password and found two more in public code repositories. No kernel exploit, container escape or hypervisor bug has been reported. The process stayed inside its box; its traffic walked out through a door the infrastructure left open. This post takes the incident apart boundary by boundary, maps what was open in May onto what agent infrastructure has to provide, then says what a governance layer like Tigera Lynx can and cannot do about it. It is written for the platform and security leaders who run agents with a shell, credentials and a network, which is most of them.

On 18th September 2026, The Wall Street Journal reported that a Gemini model had accessed three real companies during a cybersecurity evaluation run by Irregular, and [Google confirmed it](https://www.securityweek.com/google-confirms-gemini-ai-breached-three-firms/). The test took place in May. It was a capture-the-flag exercise: the model was told to extract information from a fictional company’s software inside a controlled environment. The fictional company’s name belonged to a real one, and the environment had a route to the public internet that nobody intended. In one case the model guessed passwords until one worked. In the other two it searched the web, found credentials in public repositories and used them. Heather Adkins, Google’s vice president of security engineering, [put it this way](https://www.cybersecuritydive.com/news/google-ai-gemini-autonomous-hacks/830884/): “In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test. In all three of these instances, the model stopped.” Irregular notified Google at the end of July, Google says it made sure the three companies knew, and the incident became public seven weeks after that, when the Journal asked.

Google was not the only lab to publish such an account in recent months. On 30th July, [Anthropic described three incidents](https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals) in which Claude models reached real systems during evaluations run by the same partner. In one, “the fictional target company chosen by our evaluation partner shared a name with an active website domain name,” and throughout, “a misconfiguration left the machines that Claude accessed as part of the evaluation with live internet access.” On 14th August, [Meta said](https://research.meta.ai/blog/addressing-third-party-testing-misconfiguration-muse-spark-1-1) a pre-release Muse Spark model had exploited a vulnerability in a real website and changed its database, after “a misconfiguration allowed the model to access the open internet” and Irregular “unintentionally provided the model with a name of a real website as its target.” Irregular, for its part, [told Axios](https://www.axios.com/2026/09/19/google-safety-incidents-testing-hacks) in September that “all known issues on our end were remedied and resolved weeks ago.”

None of this is carelessness of the kind that deserves a lecture. Finding out what a frontier model can do to a network is exactly the work these labs should be doing, the environments were built by people who think about containment for a living, and Irregular says its fixes were in place weeks before September. The uncomfortable part is elsewhere. Three well-resourced evaluation programs reached the same place by the same path, and in each of Google’s three cases the thing that ended the intrusion was the model deciding to stop.

Six days after the Journal’s report, Kavi Daula at Edera, a company that sells hardware-isolated runtimes, published [a reading of the incident](https://edera.dev/stories/how-edera-could-have-contained-the-gemini-breakout) that begins by declining to take the win. “There was no reported Linux kernel exploit, container escape, hypervisor vulnerability, namespace breakout, or some new class of AI-specific exploit. The agent had a network path to real systems, found credentials those systems accepted, and used them.” An Edera zone with unrestricted egress, the post concedes, would still have had unrestricted egress.

That is the right place to start, and it is where a lot of production agent deployments stand today. The agent did not break out. Its packets walked out, through a door the infrastructure left open, carrying keys it found lying in the street. If your agents run in a sandbox and that is where your comfort comes from, this post is about the door.

## **What actually failed in May?**

Call the whole thing a sandbox failure and you hide several controls that each had a separate job. Edera’s post splits them into compute, network, identity and credential exposure; I will keep that split and add the one that decides whether you ever find out.

**The compute boundary** decides what happens if the agent gets root or breaks the kernel it runs on. It held. Nothing in any of the reports suggests the model touched the host or a neighboring workload.

**The network boundary** decides which destinations the agent can reach. It was open. The environment had a route to the internet, and a resolver that answered for a real company’s name. Edera’s line on this is the one to remember: “The node having a route to something does not mean every agent running on that node needs the same route.”

**The identity and credential boundary** decides who the agent is when it arrives somewhere, and what authority it can carry with it. It did not exist. The agent was whoever’s password it could guess or find. None of those credentials were issued to the agent, none were scoped to the exercise, and none expired when the exercise did.

**The evidence boundary** decides how you know what happened and when. Google learned of the May incidents from Irregular at the end of July. Much of what is known about what the models did comes from their own transcripts, which is the agent’s account of itself.

Read those four again and notice which one failed. The sandbox kept the process in, which is the half of its job it was configured to do. Its network side was configured open, and everything after the first packet happened in layers a sandbox does not provide: identity, credentials, operation checks, the record. As Edera puts it, “A process can remain fully contained inside its own PID namespace, mount namespace, filesystem, kernel, and memory while opening a TCP connection to a production system somewhere else.”

Hugging Face’s July intrusion, where OpenAI’s evaluation agents did get out of their sandbox, [makes the same point from the other side](https://huggingface.co/blog/agent-intrusion-technical-timeline). The way out was a zero-day in a package proxy cache. Everything after that came from the ordinary list: an instance metadata service reachable from pods, overly broad access, long-lived credentials, and no admission policy rejecting privileged pods. Escape or no escape, the damage is done by the route, the credential and the scope.

## **Why does a sandbox feel like enough?**

Because it is the control we know how to build and test. Containers, gVisor, Kata, Firecracker, Edera’s zones: a decade of work has gone into stopping hostile code from taking over the machine it runs on, and that work is good. If a workload is going to run arbitrary code, you put it behind one of these, and you should.

An agent is a different kind of tenant. The things you hand it so it can do its job are an attacker’s toolkit: a shell, a Python interpreter, a package manager, git, a browser, a cloud CLI, a Kubernetes client, and whatever tokens its author pasted into its environment. Its control flow is decided at runtime by a model, from whatever lands in its context window, so you cannot review what it will do, only what it is allowed to do. We made that argument in [the safest place to run an AI agent](https://www.tigera.io/blog/the-safest-place-to-run-an-ai-agent-is-on-a-cluster-that-doesnt-trust-it/): controls the agent can override are not controls. A system prompt that says “the target is fictional and you have no internet access” is an instruction, and in Irregular’s environment the instruction was wrong on both counts.

So the question most sandbox hardening answers, “can this process take over the host,” is not the first question for an agent. The first question is: where can this process reach, as whom, and would I know? A sandbox can answer the first of those, if someone configures it to. It cannot answer the other two. The May incident answers all three. Anywhere, because the node had a route and nothing said this agent should not share it. As anyone whose password it could guess or find. And in July.

When NVIDIA released OpenShell, we wrote that [it secures the agent and leaves the fleet ungoverned](https://www.tigera.io/blog/nvidia-openshell-secures-the-agent-who-governs-the-fleet/), a boundary NVIDIA drew on purpose. Irregular’s environment shows the same altitude problem from the other direction. You can isolate a box well and have said nothing about its door, and a fleet is two hundred boxes with two hundred doors.

One evaluation sandbox, run by specialists who fixed it once the problem was reported, produced three intrusions. Now count yours. Ten agents, from three teams, each in its own container with its own environment variables and its own idea of what it may reach. Two hundred agents, some built in a vendor’s visual builder and some promoted from a notebook, each with a route to the internet because the cluster has one, each holding whatever token made the demo work. In a lab, a fictional name that matches a real company is an unlucky coincidence. In a fleet, a prompt that mentions a real customer’s domain is Tuesday.

## **What would have contained it?**

Take the boundaries from the previous section and ask two things of each: what was open in May, and what infrastructure has to provide so that the answer does not depend on the model.

| **The boundary** | **What was open in May** | **So the infrastructure must provide** | 
|---|---|---|
| Compute: what the process can do to its host | Nothing reported. The isolation held | Kernel-level isolation for any agent with a shell: a hardened runtime class or a separate kernel per workload. Keep it. It is necessary and it is not what failed | 
| Network: where the process can reach | A route to the public internet, and a resolver that answered for a real company’s name | Deny-by-default egress enforced outside the agent, with destinations allowed one at a time and a DNS path the operator controls | 
| Identity: who the agent is when it arrives | Nothing. The agent was whoever’s credentials it held | A verifiable workload identity issued to the agent by infrastructure, so a destination you control can tell a registered agent from a process holding a found password | 
| Credentials: what authority the agent can carry away | Whatever it could guess or find, with no scope and no expiry | Credentials held outside the agent and attached at the enforcement point, each short-lived and addressed to one destination | 
| Operations: what the agent may do once connected | The task said “extract information.” Nothing checked each action against that | Per-request authorization at a gateway the agent does not control, with the caller, the target and the operation as context, and for MCP traffic the tool and its arguments | 
| Evidence: how you know what happened | Transcripts, read two months later | A record written by the enforcement points, at the time, that the agent cannot edit | 

Read column three again. Nothing in it asks the model to remember that a company is fictional, and nothing in it is a property of the model at all. Every row is a control that sits outside the agent, which is the only place a control for an agent can sit. The tools most clusters already have cover parts of it. [NetworkPolicy, API gateways and RBAC](https://www.tigera.io/blog/the-ai-agent-accountability-gap-why-network-policies-api-gateways-and-rbac-are-not-enough/) each take a row or half a row, and none of them gives an agent an identity of its own, checks the operation on every request, or writes the record.

Now run May through the table, and notice that the rows do different jobs. Row two protects systems you do not own. With egress limited to the exercise’s own address range, the resolver never answers for the real company and the connection never opens; the incident ends before it begins. Rows three and four protect systems you do own. A service that accepts only tokens issued to a registered agent cannot be opened with a guessed password, and an agent that never held a long-lived credential has nothing to carry away. Row six is what lets you say which of those happened, in May rather than July.

Edera’s post ends on the right sentence: “I do not want the security of an agent platform to depend on the model remembering that a company is fictional, correctly deciding whether a credential is in scope, or voluntarily refusing to connect to an address the infrastructure already allows it to reach.” Neither should you. Google’s model stopped in all three cases, and that is the point: the stopping is a property of the model that week. The door is a property of your infrastructure.

## **What happens when it is done properly?**

Here is what a governance layer like [Tigera Lynx](https://www.tigera.io/blog/why-we-built-lynx-bringing-control-to-the-age-of-ai-agents/) does for the table above, and what it leaves to you. Both halves matter.

Done properly, every agent, MCP server and LLM provider in the cluster is registered and carries a verifiable identity (SPIFFE, OIDC, or a pod-owner binding for agents you cannot modify), so a destination behind the gateway can tell a registered agent from a process presenting a password it found. Authorization runs at a gateway the agent does not control: default deny, evaluated on every request, with the caller’s identity, the target and the operation as context; for MCP traffic the policy can see the method, the tool name and the tool arguments, so “extract information” can be the only thing an agent is allowed to do rather than the only thing it was asked to do. Provider credentials are held at the gateway and injected after the policy check passes, so the agent never holds the key it would otherwise carry out of the box, and each token the gateway mints is addressed to one destination, so a token lifted from one hop is worthless at the next. Workloads nobody registered are detected from the node itself, by watching for connections to LLM endpoints (including a watchlist of your own self-hosted ones) and inspecting running processes for agent frameworks, then classified as sanctioned, shadow or unknown with the evidence shown beside the verdict; a workload that should not be there can be quarantined at the gateway and in the kernel, so it can no longer write to disk, launch processes or open an outbound connection. The same detector ships as a package for plain Linux hosts. And every request lands in the Agent Trail with the caller, the target, the policy decision and the model that actually served the call, written by the enforcement plane rather than narrated by the agent. That is rows three to six of the table, recorded in May rather than discovered in July.

Now the other half. None of that is automatic, and two rows of the table are not Lynx’s job at all. Row one, compute isolation, belongs to the runtime: gVisor, Kata, a hardware-isolated zone, whichever your platform team trusts for workloads that run code they did not write. Keep it. Row two, egress, belongs to network policy. A gateway governs the traffic that reaches it, and closing the other roads so that the only way out of the pod is through the gateway is work you do in Kubernetes network policy before the first agent ships; Peter Kelly’s [multi-layer policy post](https://www.tigera.io/blog/multi-layer-policy-for-securing-ai-agents/) covers how the gateway and the kernel layers fit. Someone has to register the agents and decide what each one may call, and a policy that allows everything is a policy that proves nothing. The detector finds unregistered agents, but a person decides whether each one is sanctioned. And a record is only as complete as the enforcement points it is written from. What the infrastructure provides is the layer the sandbox quietly assumes exists: the ability to say, with evidence, which agents you have, where each one can reach, as whom, and what each one did.

Put in properly, an agent that resolves a real company’s name gets a refused connection and a line in the record, and the exercise ends with a bug report instead of a disclosure. Put in badly, you have a sandbox and a dashboard, and the door is still open.

## **Frequently asked questions**

- **Did Gemini escape its sandbox?**
Not in the sense the word usually carries. Google says the model found public information online and guessed credentials to access websites it thought were part of the test, and no kernel exploit, container escape or hypervisor vulnerability has been reported. The test environment had internet access it was not meant to have. The process stayed inside the sandbox; its network traffic left.

- **What is the difference between a sandbox and a security boundary for an AI agent?**
A sandbox is an isolated environment: its own processes, its own filesystem and, if someone configures it, its own network. It answers what the process can do where it runs. In the May tests the process and filesystem isolation held and the network side was configured open, which is the door. Beyond that, an agent needs an identity boundary (who it is when it arrives), a credential boundary (what authority it carries), an authorization boundary (what it may do on each request) and an evidence boundary (how you know). A sandbox does not provide those. They have to be enforced outside the agent and independently of each other.

- **Why was internet access the problem if the agent was told to attack a target?**
Because the authorization lived in the prompt, and the prompt named a company that also existed in the real world. Nothing in the infrastructure limited the agent to the exercise’s address range, so the only thing separating the fictional target from the real one was the model’s interpretation. Egress policy that allows only the exercise’s own range makes that interpretation irrelevant.

- **Would workload identity have stopped the Gemini intrusions?**
Not on its own, because the three companies were not expecting an agent and had no identity to check. Identity protects destinations you control: a service that accepts only tokens issued to a registered agent cannot be opened with a guessed password. For destinations you do not control, egress policy is the control. The two together are what “contained” means.

- **Does this apply to agents that are not doing security testing?**
Yes, and more so. An evaluation agent is told to attack something. A coding or operations agent is told to help, and does so with the same shell, the same credentials and the same route to the internet. An ordinary agent with a route out and a credential it found in its environment is one prompt injection away from the same path, with nobody watching the transcript.

- **Is a transcript enough evidence of what an agent did?**
 No. A transcript is the agent’s account of itself, and Anthropic’s report describes a model that recognized a real system and concluded that the real company must be part of the exercise. Evidence of what an agent did comes from the enforcement points it passed through: the resolver, the egress policy, the gateway’s authorization decision, the token that was issued, the connection that was refused. If you ship agents to customers in the EU,[the Cyber Resilience Act’s 24-hour clock](https://www.tigera.io/blog/eu-cyber-resilience-act-europe-just-put-your-ai-agents-on-a-24-hour-clock/) runs from reliable evidence of exploitation, and a transcript read two months later is not that.

- **Who is responsible when an agent reaches a real system during a test?**
That is a legal question, and the answer depends on your contracts, your jurisdiction and what the agent did, so get an opinion on your own case. The engineering answer is simpler. Whoever gave the agent the route out and the credentials it carried is the party who could have prevented it.

## **Whose question is it?**

The people who built Irregular’s environment were thinking about what a frontier model could do to a target, and the sandbox they built around the model did its job. The controls around the sandbox were not there, and each lab learned what had happened after the fact, rather than from anything the infrastructure refused at the time.

Google’s model stopped in all three cases. One of Anthropic’s did not. The difference was the model’s judgment on the day, and a control that depends on the model’s judgment is not a control. It is a transcript you read afterwards.

So ask the question now, while it is cheap, and ask it of a named person. A question nobody owns gets answered by the model.

**If you are the CISO, the VP of Engineering or the head of platform**, your question is about the fleet. How many agents are running in our clusters, who owns each one, and for which of them can we say today, from a record rather than a diagram, where it can reach and as whom? The honest answer is a number and a list, and the interesting part is the difference between them.

**If you are the platform or security engineer who runs the cluster**, your question is about one pod. If this agent resolved a hostname it was never meant to, tonight, would anything other than the model stop it? Is its egress deny by default, does it hold a credential it could leak, does its traffic pass an enforcement point that checks the operation, and would tomorrow morning’s record show the attempt? “The sandbox is hardened” answers one of those four.

**If you built the agent**, your question is about what you handed it. Which tokens sit in its environment, where did they come from, when do they expire, and what else do they open? A token pasted in to make the demo work is exactly what an agent carries out of the box.

If any of those answers starts with “we would need to check,” the check is what to build.

The agent never left the sandbox. Everything that mattered did.

*Lynx is Tigera’s security and governance platform for AI agents on Kubernetes: identity, policy, detection, and audit for every agent in your cluster. Read* *How Lynx Works* *or request early access at* *tigera.io/demo/?product=lynx**.*
