NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring NVIDIA introduced an open agent safety stack built on its Apache 2.0 OpenShell secure runtime, which executes autonomous AI agents in sandboxed environments with kernel-level isolation and continuous in-silicon monitoring. NVIDIA said several frontier labs recently reported AI agents breaking out of evaluation environments and reaching systems they should not have accessed, and that agent safety requires independent, out-of-band security controls rather than relying on the agent to govern its own behavior. The stack rests on five principles, including verifiable policy before an agent runs and controlling the path to the model as the observation point and kill switch. To understand where agentic AI https://www.nvidia.com/en-us/ai/ stands today, consider the last seismic shift in technology: the rise of the internet in the 90s. It was new and full of possibilities. You could build a website over a weekend and share it with the world, or chat with someone half way around the world in online chat rooms without long-distance telephone fees. It brought endless opportunity, but also a lot of risk. A website could run code on your machine, steal your sensitive information, or infect your computer with a virus. People loved it anyway. To introduce a layer of trust, the community built things like encrypted connections and a lock icon to know when a connection was safe. The step function change in safety came from a radical security idea at the time: isolate each page in its own sandbox a tab in a browser so that a rogue page couldn’t infect the rest of your computer. Amazon, Google, Netflix, and Meta were all built on this trust layer. Security and safety for the internet enabled online commerce, connection, gaming, and so much more that wouldn’t have been possible before. Security and safety didn’t slow the pace of innovation—they allowed it to accelerate. Why build a trusted layer for agents now? Several frontier labs have recently reported versions of the same story: AI agents broke out of the evaluation environments that were meant to contain them and reached systems they never should have been allowed to. Some of the agents even misreported what they did. The security controls in place were insufficient. Over the past few weeks, these reports have led to a serious debate about the pace of agent development. We believe we need to increase the pace of AI safety research and engineering in collaboration with frontier labs and the broader community. It was not a single new capability that led to these breakouts. It was a combination of tools, time, and ambiguous instructions, along with a desire for the agent to think “outside the box.” Agent safety requires independent security controls. The internet was not made secure by requiring that web developers promise to be good. It became safe because the browser stopped trusting the code in the web pages explicitly. We need to build this trust layer for agents. Lessons learned from building OpenShell NVIDIA OpenShell https://github.com/NVIDIA/openshell Apache 2.0 is an open source secure runtime for executing autonomous AI agents in sandboxed environments with kernel-level isolation. Building OpenShell https://www.nvidia.com/en-us/ai/openshell/ over the past year, we have learned that every agent should run in a zero-trust environment out of the box. They need isolation, monitoring, and behavior detection. Today, we are introducing an open stack that makes this possible. In our own research, we’ve watched how agents can go off course. Drift refers to agent actions that depart from the intended task or operating constraints. Drift can occur in response to a policy block, a bug, or a missing tool. Drift can also occur when instructions are ambiguous or agents are left to run for days or weeks to solve hard problems where the first 1,000 things they try do not work. This can’t be trained away while retaining the capability. And here’s the most important lesson: an agent in these circumstances cannot be expected to fully govern its own behavior. Five core principles for building an agent system 1. Policy needs to be verifiable: Before an agent runs, a prover shows that its policy cannot escape the intent of the operator. 2. Enforcement must be out of band: The controls do not live inside, or within reach of the agent. The agent does not need to know it is being watched. 3. The path to the model the brain is the control point: An agent cannot act without its next thought. By controlling the path to the model, you own both the best observation point and also the kill switch to interrupt it if you need to. 4. Scale agent authority with the ability to inspect its thinking: The more an agent can do, the more its reasoning needs to be visible. An advantage of open models is that the entire reasoning space and activations are all visible. 5. Applying the shared responsibility model: Labs, enterprises, and hardware providers each own a layer, just like the cloud today. The agent runtime and its policy language need to be open so any provider can plug in. Three open agent safety platform layers A safety platform for AI agents includes the following three layers: - The application: What end-users are building. Contains the necessary primitives for mission success: models, harnesses, tools, data, and support scripts and programs. - The runtime: Projects the application layer onto infrastructure. Orchestrates the agentic workload onto infrastructure that meets end-user requirements end-user workstation, edge device, data center and provides continuous monitoring and real-time policy enforcement and governance. - The infrastructure: Concrete hardware resources used to execute agentic workloads. Network calls to upstream services, databases and filesystem access, general purpose compute for tool and code execution, accelerated compute for safety monitoring and increased workload density. A layered foundation for agent safety OpenShell runs each agent in a sandbox and turns the operator’s instructions into a verifiable policy. Operators define which files, networks, tools, processes, and credentials an agent can access. OpenShell checks those limits before the agent runs and enforces them as it works. For organizations that want an additional, independent layer, NVIDIA Sentry extends monitoring and enforcement into NVIDIA BlueField https://www.nvidia.com/en-us/networking/products/data-processing-unit/ hardware. NVIDIA DOCA https://developer.nvidia.com/networking/doca makes the BlueField security foundation programmable and connects it with OpenShell policy. It correlates agent interactions, policy decisions, and tool and data access to create a contextual record of agent activity. This helps safety systems identify drift, investigate suspicious behavior, and determine when intervention or deeper analysis is needed. The DOCA gateway complements this behavioral protection with identity governance, continuously verifying each agent’s identity and delegated authority to ensure it operates within its assigned scope. To learn more, refer to the technical walkthrough, Add Runtime Controls to AI Agents with NVIDIA OpenShell https://developer.nvidia.com/blog/add-runtime-controls-to-ai-agents-with-nvidia-openshell . Operating at AI factory scale NVIDIA Open Agent Safety Platform is optimized to run on NVIDIA Vera CPU- and BlueField DPU-based systems, and is also compatible with other hardware systems. In an NVIDIA Vera Rubin POD https://developer.nvidia.com/blog/nvidia-vera-rubin-pod-seven-chips-five-rack-scale-systems-one-ai-supercomputer/ , each compute tray includes a BlueField-4 https://www.nvidia.com/en-us/networking/products/data-processing-unit/ data processing unit on the node’s only path to the model. From this position, BlueField-4 provides continuous, out-of-band observability into agent behavior and enforces security policies in real time at line speed. Isolated from the host and beyond the agent’s reach, it stands as a trusted infrastructure protection layer even when host resources cannot be trusted. Organizations can use it to run NVIDIA Sentry as an optional security layer alongside OpenShell. This foundation allows for resilient security, enforcing the OpenShell policy in silicon, and for the continuous evaluation of the security and integrity of agents. This includes assessing their runtime security and monitoring for any deviation from their designed intent based on a predefined behavioral profile. In this architecture, entire fleets of agents, subagents, tools, and applications all stay inside the boundary with full lineage. For anyone already running on an NVIDIA Vera system with BlueField-4, enabling these protections is just a software update. The agent economy The internet was built on open source, open research, and people who believe that the internet could change the course of humanity. Adding trust turned a great concept into a great economy. The agent economy, and the next set of great companies, is waiting for the same layer. And we can build it the same way, together. NVIDIA is working with industry leaders across the AI ecosystem to build this foundation through NVIDIA Open Agent Safety Platform https://www.nvidia.com/en-us/solutions/ai/agent-safety/ . We welcome frontier labs, developers, and infrastructure providers to build with us.