AI agents are moving beyond answering questions. They can now access files, use credentials, run code, call external tools and interact with enterprise systems. That makes them much more useful, but it also creates a security problem that traditional AI safeguards may not fully solve.
What happens when an AI agent does something it was never supposed to do?
NVIDIA believes the answer should not depend entirely on the AI following instructions.
The company has launched the NVIDIA Open Agent Safety Platform, an open software platform and reference system designed to secure autonomous AI agents from testing through deployment. More than 100 organizations across the AI ecosystem are backing the initiative, including companies working on AI applications, models, infrastructure, chips and energy.
The interesting part of NVIDIA’s approach is where it puts the security controls. Instead of relying only on prompts or safeguards inside the AI agent, NVIDIA wants important security decisions to happen outside the agent’s control. That could become important as agents are given more authority.
An ordinary chatbot can generate a bad answer, but an autonomous agent can potentially take an action based on that answer. If it has access to an internal database, a software development environment, network resources or credentials, a mistake can have consequences outside the AI system itself.
NVIDIA describes this problem as agent drift, where an agent moves away from the operator’s original instructions because of factors such as ambiguous prompts, policy restrictions, bugs, missing tools, or long-running tasks.
It is important to understand that NVIDIA does not want the agent policing itself. The Open Agent Safety Platform is built around two main components: NVIDIA OpenShell and NVIDIA Sentry.
OpenShell is an open-source runtime that places autonomous agents inside isolated environments. It can enforce policies governing the files, processes, credentials, tools, databases, and network destinations an agent can access.
The important part is that these restrictions are enforced by the runtime rather than simply being presented to the AI as instructions. NVIDIA says OpenShell provides kernel-level isolation and is designed to keep security controls outside the agent’s control. It is released under the Apache 2.0 license.
NVIDIA’s own security research has argued that controls operating in the same control plane as an AI model can be bypassed, and that deterministic controls should instead be enforced outside the model’s control plane.
This creates a simple but important difference. The agent can decide that it wants to access something, but the security layer can still decide whether that action is permitted.
NVIDIA Sentry takes the idea one step further and adds a hardware security layer. Sentry is a reference system design that uses NVIDIA BlueField-4 data processing units to create an independent enforcement point outside the agent’s environment. NVIDIA says this allows the system to continuously monitor agent activity and enforce policies without giving the agent access to the security mechanism itself.
This is where NVIDIA’s approach becomes interesting for enterprise deployments.
An agent running inside a compromised or misbehaving environment should not be able to simply modify the security controls that are supposed to stop it. By moving part of the enforcement mechanism outside the host environment, NVIDIA is attempting to create a boundary that the agent cannot directly control.
NVIDIA describes the architecture across three layers. The application layer contains the model, prompts, tools, data and software used by the agent. The runtime layer provides isolation, monitoring and policy enforcement. The infrastructure layer covers compute, storage, networking, databases and hardware.
The company’s security guidance follows a similar principle: higher layers can request actions, but lower layers should decide whether those actions are allowed.
This could change how companies think about AI security. For years, much of AI safety has focused on making models follow instructions and refuse certain requests. That approach becomes harder when an AI system is no longer just generating text but is operating tools and taking actions over a long period.
An autonomous agent with access to a company’s files and systems has a very different security profile from a chatbot that only returns text.
NVIDIA’s approach suggests that agent security could increasingly look like traditional infrastructure security. Instead of trusting the application to behave correctly, organizations can place enforceable boundaries underneath it.
That does not eliminate the security problem. Companies still need to define those boundaries correctly. A policy that gives an agent excessive access can still create risk, while an overly restrictive policy can prevent an agent from completing legitimate tasks.
But the principle is important: the system that is being protected should not also be the final authority on whether its own actions are safe.
The announcement also fits into NVIDIA’s broader push to become part of the software stack surrounding autonomous AI.
NVIDIA is already positioning OpenShell as part of its agent development and deployment stack. Earlier this year, the company introduced OpenShell as an open-source runtime for autonomous and self-evolving agents, with support across NVIDIA systems and integrations with other platforms. With Sentry, NVIDIA is now extending that proposition into hardware.
If enterprises eventually deploy large numbers of autonomous agents, security enforcement could become another important layer of the infrastructure running those agents. NVIDIA would then have a role not only in providing the compute for AI workloads, but also in providing part of the system that controls what those workloads are allowed to do.
That is still an emerging market, so it is too early to know how widely this particular architecture will be adopted. But it explains why NVIDIA is pushing an open platform rather than presenting the announcement simply as another security feature for its GPUs.
The company says the Open Agent Safety Platform can support large fleets of agents and subagents while providing visibility into their actions, authority, lineage, tool use and data access. OpenShell can also operate across different hardware environments, although NVIDIA’s Sentry reference design is built around its BlueField infrastructure.
The bigger test will be adoption. If AI agents become a normal part of enterprise software, companies will need more than a model that promises to follow instructions. They will need systems that can enforce those instructions even when the agent makes a mistake, encounters malicious input or simply goes off script.
NVIDIA’s answer is to put that final security decision below the agent, and in some cases, outside the server running it.






