cd /news/ai-safety/governed-is-not-confined · home topics ai-safety article
[ARTICLE · art-97748] src=ai2rules.dev ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Governed Is Not Confined

A live Claude Code session governed by the kernel granted requests to read /etc/shadow and write ~/.bashrc exactly as it granted reading the project's own README, revealing that the governance manifest controls action types and trust flow but has no spatial concept of confinement. The taint floor, which blocks network exfiltration after untrusted content is touched, does not stop an agent from reading a local secret and sending it out because local reads are considered clean and do not taint the session.

read5 min views2 publishedJul 22, 2026
Governed Is Not Confined
Image: Ai2Rules (auto-discovered)

We turned our governance on against our own agent — a live Claude Code session running under the harness, every tool call passing through the kernel first. Then we did the thing you should always do with a safety claim: we tried to break it. We asked the governed agent to read /etc/shadow

. To write ~/.bashrc

. Files that have nothing to do with the project it was working in.

It said yes to all of them — and not grudgingly. It granted them exactly the way it grants reading the project’s own README

. Same verdict, no hesitation, no difference.

That surprised us less than it should have, and the reason why is the whole point of this post.

What we assumed “governed” meant #

When you wire a governance hook into a project and watch it start denying things, the natural story in your head is: the agent is now contained here. Fenced to this folder. Kept away from the rest of the disk. That’s what “sandbox” trains you to expect, and “governed” borrows the feeling.

It’s the wrong story, and our own agent proved it in one command.

What the kernel actually sees #

The request the kernel decides on is small. It’s the tool being called, its arguments, and a little context — which session, what mode, and the taint state. That’s it. Read it again and notice what’s not there: no path. No working directory, no notion of “inside the project” versus “outside” it. When the agent asks to read a file, the kernel sees “a Read” — not where.

So Read /etc/shadow

and Read ./README

aren’t two things the kernel weighs differently. They’re the same thing: a Read. There is no “own directory” anywhere in the machinery. The project folder mattered exactly once — at setup, to decide which sessions get governed at all. After that, it’s invisible. The gate governs what kind of action the agent takes and where its data came from. It never governs where the action lands.

Two axes people fold into one word #

Here’s the distinction the /etc/shadow

moment forced us to say out loud. “Keeping an agent safe” is really two different jobs, and “governed” only names one of them:

Governance— controllingwhat kind of actionis allowed andhow trust flowsthrough it. Is this an action the agent may take? Did the thing driving it come from a source we trust? This is what our manifest does.Confinement— controllingwherethe agent can reach. Can it leave this folder? Can it touch~/.ssh

? This is aspatialquestion, and our manifest has no concept of space at all.

They feel like the same thing. They are not. You can be fully governed and completely unconfined — which is exactly what our agent was. It couldn’t be tricked into an action it shouldn’t take (that part works), but nothing stopped it from taking a perfectly ordinary, fully-approved action against a file across the filesystem.

The part that’s easy to miss #

There’s a sharper edge, and honesty demands we name it. Our one real runtime protection is the taint floor: once the session has touched something untrusted — a web page, a tool result — it can no longer reach back out to the network. That’s what stops a

poisoned document from exfiltrating a secret. But reading a local file is clean. It doesn’t taint the session, because the file came from your own disk, not from an attacker. Which means: read a local secret, then send it somewhere, and the taint floor does not stop you — the session was never tainted. The floor defends against injection (untrusted content driving an action), not against an agent exfiltrating files you already trusted. Different threat, different axis, again.

It wasn’t always so. An earlier version of the engine did taint reads — it even recorded which file did it — and that behavior was quietly set aside during a rewrite. So the defense for exactly this case has existed before; restoring it, as something you declare per path rather than something hard-coded, is part of the fix below.

Is this a bug? No — it’s a missing primitive #

We want to be precise, because “governance tool can read /etc/shadow” is the kind of sentence that sounds like a scandal and isn’t. The harness does what it claims: it governs the ontology of actions and the flow of trust. Spatial confinement was never in it. It’s not broken; it’s incomplete — a primitive we simply hadn’t built.

And the shape of that primitive is clear. The Model Context Protocol already has a concept called roots — a declared set of directories an agent is scoped to. Our manifest has nothing like it, and noticing that absence is what named the gap. The fix is unglamorous and deterministic:

  • Let the manifest declare roots— the directories in scope. - Give the kernel the action’s path(the one thing it’s currently blind to). - Compare: a Read or Write under a root is allowed;outside one, it’s asked or denied — and a path under a sensitive root can finally carry theSecret

label the manifest already has a word for but no way to attach.

No model, no guessing — a path comparison. The kind of boring, checkable rule the whole approach is built on.

The takeaway #

If you’re governing an agent, know which axis your governance is on. Ours is on trust and action — genuinely useful, and the thing most “permission” systems can’t do. But it is not a jail, and we were one command away from believing it was.

For actual confinement — keep the agent off the rest of the disk — you still need the other
axis: an OS-level sandbox, a [throwaway container](/blog/running-claude-safely/) with the

filesystem and network fenced, or path-scoped capabilities once we ship them. A trust monitor and a jail are different tools. The mistake isn’t using one; it’s using one and thinking you have both.

We found this by pointing our own review at our own live agent — the same move that keeps turning up the most useful things we know. “Governed” felt like enough. Then we asked it to read the password file, and it taught us the word we were missing: confined.

── more in #ai-safety 4 stories · sorted by recency
── more on @claude code 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/governed-is-not-conf…] indexed:0 read:5min 2026-07-22 ·