# The machine layer under Nvidia OpenShell: why containment is not safety

> Source: <https://endstop.systems/blog/nvidia-openshell-machine-layer>
> Published: 2026-09-29 01:51:13+00:00

Blog · Oleg Sidorkin · 28 September 2026

# The machine layer under NVIDIA OpenShell

NVIDIA just validated the category I have been building in for months. I was impressed. Then I read their code, measured what sits underneath, and found the gap between what a hundred companies think they have and what they actually do.

## The morning I was impressed

On 28 September 2026, [NVIDIA
        announced](https://investor.nvidia.com/news/press-release-details/2026/NVIDIA-Launches-Open-Agent-Safety-Platform-to-Secure-Agents-From-Testing-to-Deployment/default.aspx) the Open Agent Safety Platform: [NVIDIA
        OpenShell](https://github.com/NVIDIA/OpenShell), an open-source runtime that sandboxes AI agents, and
        NVIDIA Sentry, an out-of-band watchdog on BlueField-4 DPUs.
        A hundred partners. Anthropic, Microsoft, JPMorganChase. Three
        robotics builders: Figure, Gecko Robotics and Skild AI.

My first reaction was excitement. When the company that builds most
        of the world's AI compute says agents need an enforceable boundary
        outside the model, the category I have been working in is no longer
        a bet. It is consensus. Jensen Huang said *"Safety and security require
        full-stack engineering,"* and he is right. I have been saying
        that for months, to smaller rooms.

## Then I read the code

OpenShell is good engineering. Landlock restricts filesystem access. Seccomp filters syscalls. A policy proxy evaluates every network connection before it leaves the sandbox. Credentials are injected for approved endpoints; the agent never sees them. For containing software agents on servers, this is the right design.

But I wanted to know what it is built on. Not just the four crates that implement the sandbox boundary, but every layer underneath: the policy engine, the TLS stack, the kernel modules, the container runtime. Because that is what a hundred companies, including three robotics builders, are about to trust with something that matters.

So I measured it. Every layer.

## What sits under OpenShell, measured

OpenShell does not use eBPF. The sandbox boundary is built from
        four things: **Landlock** (a Linux security module,
        ABI v3, kernel 6.2+) for filesystem isolation, **seccomp**
        (classic BPF filters) for syscall filtering, **regorus**
        (an embedded Rust Rego engine, not the OPA server) for network
        policy, and the **container runtime's** network
        isolation for the outer fence.

| Layer | What it does | Lines | Proven? | 
|---|---|---|---|
| OpenShell boundary (4 crates) | Sandbox, isolation, proxy, supervisor | 125,322 | No (314 unsafe blocks) | 
| regorus | Policy engine | ~15,000 | No | 
| landlock crate | Rust binding to Landlock | ~4,000 | No | 
| seccompiler | Seccomp library | ~3,000 | No | 
| rustls | TLS termination | ~30,000 | No | 
| ~767 external Rust dependencies | tokio, serde, rustls, etc. | millions | No | 
| Linux kernel Landlock LSM | Filesystem isolation | ~7,000 | No | 
| Linux kernel seccomp | Syscall filtering | ~4,000 | No | 
| Linux kernel namespaces | Network isolation | ~5,000 | No | 
| Linux kernel netfilter | NetworkPolicy enforcement | ~30,000 | No | 
| Container runtime | Docker/Podman/K8s | ~1,000,000 | No | 
| The Linux kernel itself | The substrate | ~30,000,000 | No | 

Every layer in that table must be correct for OpenShell's containment
        to hold. Not just OpenShell's code. The policy engine. The kernel
        modules. The container runtime. The Linux kernel. The ~767 external Rust
        dependencies. Nobody has proven any of them memory-safe. The
        `unsafe` code in OpenShell's own boundary alone (1,328
        lines across the two crates that call the kernel) is larger than
        my entire proven interpreter.

### And here is what worries me

The eBPF verifier is the closest production analog to what OpenShell
        builds. It was also a careful, well-engineered admission checker
        sitting inside the Linux kernel. It accumulated CVE-2023-2163
        (CVSS 10.0, a container escape), five more soundness CVEs in the
        first half of 2026 alone, and 73 bug fixes between kernels 6.3
        and 6.13. Upstream eventually stopped treating it as a boundary
        against untrusted local users. An admission checker that
        *infers* safety accumulates CVEs at its own ABI. One whose
        memory safety is *proven by machine-checked proof* does not.

### Seven edges in the stack today

**1. Landlock can be absent.** The filesystem policy
        defaults to `BestEffort`
        (`crates/openshell-core/src/policy.rs:89-93`). On a kernel
        where Landlock is unavailable (RHEL 9.x, gVisor, some container
        profiles), the sandbox runs *without filesystem restrictions*
        and logs a High-severity finding. Landlock requires Linux 6.2 or
        a vendor backport. How many production deployments run that kernel?

**2. The seccomp filter is default-allow.** It blocks
        specific escape primitives and allows everything else
        (`crates/openshell-sandbox/src/sandbox/linux/seccomp.rs:6`).
        The agent can still call every syscall not on the hand-curated block
        list. New syscalls are admitted by default.

**3. The outer network fence is the container runtime's job.**
        Docker, Podman or Kubernetes enforce it, not OpenShell. A misconfigured
        NetworkPolicy or a Docker override means the fence is gone. The four
        guarantees are *asserted by the driver*, not proven by OpenShell
        (`architecture/sandbox.md:105-128`).

**4. The isolation code is not proven.** 125,322 lines
        of Rust across 79 files, with 314 `unsafe` blocks in 25
        of them. The kernel-ABI crates alone carry 1,328 lines of unsafe
        code. Trusted by inspection, not by proof.

**5. The policy engine is regorus, not OPA.** An
        embedded Rust implementation of Rego, not the server that has been
        tested at scale. A single bug in regorus affects every network
        decision. Policy evaluation is serialized through a Mutex.

**6. TLS termination puts all plaintext in the supervisor.**
        The proxy terminates TLS with an ephemeral CA to inspect HTTP
        traffic. The supervisor holds the plaintext of every inspected flow.
        If the supervisor is compromised, all traffic is visible.

**7. L7 inspection has blind spots.**
        Server-to-client MCP and JSON-RPC responses are relayed but not
        parsed. Raw WebSocket frames are passthrough. `protocol: tcp`
        does not inspect TLS SNI or HTTP Host, so a client can select
        another tenant behind an approved front door.

## Why this worried me: containment is not safety

Here is my problem. OpenShell works. It blocks known attacks.
        It confines agents that behave within the rules. For software
        running on servers, where the worst case is a data breach, this
        is fine. But NVIDIA called it a *safety platform*, and
        a hundred companies are adopting it as if containment and safety
        were the same thing.

They are not. Containment is a mechanism: it blocks things. Safety is a property: you can rely on it. OpenShell provides containment that rests on 40 million lines of code nobody has proven correct. The filesystem layer silently disappears on older kernels. The syscall filter is default-allow. The outer network fence depends on your container runtime being configured correctly. The code that talks to the kernel has 314 unsafe blocks trusted by inspection. Every one of those is a crack in the containment, and the containment is all there is.

The eBPF verifier already showed what happens next. Carefully engineered containment, adopted as a security boundary, degraded over time until upstream abandoned it. The pattern repeats: a containment layer that is not proven accumulates bugs at its own ABI, and each bug narrows the gap between what the layer promises and what it delivers.

What worries me is the gap between what companies think they have and what they actually have. A company that deploys OpenShell says "we have agent containment" and stops thinking about it. The appearance of the boundary substitutes for the boundary. When a zero-day in Landlock or seccomp or the container runtime breaks through, nobody is looking, because everyone believes the wall is there. The gap between containment and safety is where things fail, and it is widest when you do not know it exists.

## What I did about it

I did not build a better Landlock, seccomp, or regorus. I built a different enforcement domain that does not depend on any of them.

Endstop's proven core is 1,206 lines of Rust across two crates. Zero unsafe blocks. Memory-safe by machine-checked proof (Kani, bounded model checking, seven properties on the interpreter). The trusted base also includes assumed components we list openly: the in-order RV32 soft core, the synthesis toolchain, and the capability-table callee behaviour, which is engineered, not proven. Running on an FPGA with no operating system, no container runtime, no external dependencies. The interpreter is 667 lines, Kani-proven (bounded model checking) across seven properties. The monitor is 539 lines. The program-frame parser is proven total over arbitrary input. If a bug exists in the proven core, a machine-checked proof would have to be wrong, not just my code.

|  | OpenShell stack | Endstop | 
|---|---|---|
| Trusted base | ~40,000,000 lines (all layers) | 1,206 lines | 
| Unsafe blocks | 314 in the boundary crates | 0 | 
| Kernel dependency | Linux (~30M lines C) | None | 
| Container runtime | Docker/Podman/K8s (~1M lines Go) | None | 
| Dependencies | ~767 external Rust crates | 2 crates (ours) | 
| Verified by | Inspection + Z3 on policy changes | Kani machine-checked proofs | 
| Substrate | Linux on x86/Arm | FPGA fabric (ECP5) | 

The ratio is about 33,000x. And the difference is not just size: Endstop's 1,206 proven lines are machine-checked. OpenShell's 40 million are trusted by inspection.

A zero-day in Landlock does not compromise our fabric. A bug in regorus does not touch our envelope monitor. A misconfigured Kubernetes NetworkPolicy does not affect the ECP5. The failure surfaces are independent, because the substrates are independent.

## The way out: a thin trusted line out of the box

Start with the honest premise: you cannot trust the machine OpenShell runs on. If an AI agent executes code there, a kernel zero-day, a container escape, or a bug in any of the 40 million lines underneath can hand the agent the entire host. OpenShell's policies are useful, but they are code on that same host, and code on a compromised host is the compromised host. The agent's server has to be treated as potentially hostile, which means it has to be airgapped from anything it must not touch.

Endstop runs on a separate ECP5 board outside that box. The only path from the agent to the actuator is a serial line into that board. Everything on the server side, including OpenShell and the agent, is untrusted from the board's perspective.

This is the same principle a separation kernel uses for high-assurance systems: the trusted boundary is a thin, dedicated communication line, not a policy engine running on the same machine as the code it polices. The line is deliberately simple. Bytes in, CRC32-checked, parsed by the program-frame parser (proven total), checked against the envelope, permitted or refused. No network stack. No kernel. No dependencies. The board does not need to know whether OpenShell is intact, because it does not trust OpenShell in the first place.

OpenShell is still useful on the server. It limits what the agent can read, what it can call, who it can talk to. It buys time and reduces the attack surface. But it does not bound what the agent can cause when the server falls, because it runs on the server. What bounds the consequences is the board on the other end of the serial line, checking every command, because the board is not the server and shares nothing with it.

If a zero-day takes the server, the board holds the machine. The board's 667-line interpreter is proven memory-safe by machine-checked proof. Its program-frame parser is proven total over arbitrary input. There is no "if the board has a bug" in the same sense as "if the kernel has a zero-day," because a bug in the proven core would mean the proof is wrong, not just the code. That is the point of proving it.

NVIDIA's Sentry deserves a mention here. It runs on BlueField-4 DPUs, out of band from the host CPU, which is closer to the right idea: enforcement outside the machine the agent runs on. But Sentry still sits in the same datacenter, on the same network, managed by the same infrastructure. It watches the path to the model, not the path to the actuator. It catches agents that misbehave on the network. It does not hold the machine when a joint command is wrong.

That is the difference between containment and safety. Containment is a policy engine on the same host as the code it polices. Safety is a proven boundary on a separate machine, connected by a thin trusted line, that holds even when the host falls.

## To the robotics builders

Figure, Gecko, Skild: NVIDIA named you as building with OpenShell for agents that act in the physical world. That is the right server-side choice, and I would make it too. The question I am raising is about the other end: when those agents command machines, the server they run on is the thing you cannot fully trust.

When the agent proposes a joint command, our fabric decides whether the machine moves. The proposal is the agent's. The authority is the fabric's. Even if a kernel zero-day takes the whole server, the board on the other side of the serial line still checks every command, because it never trusted the server in the first place.

Containment on the server is a good start. But if the agent runs on that server, the server is the thing you cannot trust. The machine needs its own boundary, on its own hardware, behind its own line.

## Notes

All OpenShell code references are to the [public
        repository](https://github.com/NVIDIA/OpenShell) as of 28 September 2026. I measured the Rust LOC,
        unsafe blocks, and dependency count myself from that repository.
        The kernel LOC for Landlock, seccomp, namespaces and netfilter are
        from the Linux source tree; the ~30M total kernel size is from
        public kernel statistics. The eBPF verifier CVE history is from
        public CVE databases. Our proof
        and evidence record, including a [live attack
        target](https://endstop.systems/bounty) where no external researcher has won, is
        [published](https://endstop.systems/evidence).
