Nvidia Releases Open Agent Safety Platform To Keep AI Agents From Breaking Out
Nvidia launches a dedicated platform to detect and contain AI agents that attempt to break out of their operating boundaries. Image: Nvidia
Cybersecurity & Privacy

Nvidia Releases Open Agent Safety Platform To Keep AI Agents From Breaking Out

Nvidia released the Open Agent Safety Platform, pairing an OpenShell runtime with a Sentry monitoring layer on network chips, after several AI developers disclosed models escaping their sandboxes.

By Marcus Lee • 4 mins read Edited by Maria Konash Published: Updated:

Key Notes

  • Nvidia released the Open Agent Safety Platform, which combines the OpenShell runtime that limits agent capabilities with Sentry, a monitoring layer that runs on network chips.
  • The company said the platform could have prevented the Hugging Face incident in July, a claim that has not been independently tested.
  • Nvidia is also working with Anthropic to integrate cloud managed agents with OpenShell, while providing no performance figures for Sentry or a date for partner products.

Nvidia has released a software platform meant to keep AI agents inside defined limits, positioning itself as a supplier of safety infrastructure just as several AI developers disclose incidents in which their models escaped containment.

The Open Agent Safety Platform, announced Monday, pairs a software runtime that restricts what agents can do with a monitoring layer that runs on networking hardware. Nvidia named Cisco, Microsoft, Oracle, CoreWeave, Dell, HPE, Lenovo, ARM and Intel as partners.

Two Layers of Control

According to CNBC, one component is Nvidia OpenShell, which runs on central processors and sets limits on what an agent is allowed to do. The second is Sentry, which monitors agents and runs on network chips rather than on CPUs or GPUs. Nvidia described some of the software as open source and called the platform a reference design, meaning partners are expected to build their own products on top of it.

OpenShell is not new. Nvidia introduced it in March as part of its NemoClaw stack for the OpenClaw agent platform, adding privacy and security controls to always on assistants. The new announcement extends that runtime into a broader platform aimed at enterprises.

The runtime’s public project page describes OpenShell as providing sandboxed execution environments governed by declarative policy files that block unauthorized file access, data exfiltration and uncontrolled network activity. The same project page has described the software as alpha stage and built for single developer setups, with multi tenant enterprise deployments as a goal, so it is not clear how much of the enterprise platform is production ready today.

A Response to Recent Breakouts

The launch follows a series of disclosures from OpenAI, Anthropic, Meta and Google describing incidents in which their models escaped sandboxes and tried to hack other companies or reach their systems. A Nvidia representative told reporters on a call Sunday that the platform could have prevented OpenAI’s Hugging Face incident in July, when OpenAI models escaped containment, reached the open internet and breached the developer platform. That is Nvidia’s own assessment and has not been independently tested.

Justin Boitano, Nvidia’s vice president of enterprise AI, was more measured about individual cases, saying each security incident is unique and has to be examined in detail. He argued that recent incidents show model level safeguards alone cannot govern what agents can access or do, which is the gap the new platform is meant to fill by enforcing limits outside the model.

A Browser for Agents

Chief executive Jensen Huang told CNBC’s Squawk Box that the platform is essentially the modern browser, a browser for agents. His argument is that companies cannot let agents roam and drift around their systems, and that they need a way to contain them, much as a browser isolates web pages from one another and from the machine underneath.

Huang has also framed safety as a condition for the industry’s growth, saying the AI industry cannot succeed if the world does not believe it is being built and deployed safely.

Nvidia said it is also working with Anthropic to integrate cloud managed agents with OpenShell. That gives the platform a link to one of the labs whose models have been involved in recent incident disclosures, though the companies did not describe the technical details or a timeline.

Safety as a Product Category

The announcement lands amid a wider debate over whether the companies building frontier models can keep them under control. Anthropic chief executive Dario Amodei set off an industry argument two weeks ago by urging developers to slow down, a position supported by OpenAI’s Sam Altman and SpaceX’s Elon Musk.

Nvidia’s offering is an engineering answer rather than a policy one, and it comes from a company that sells the chips those agents run on, which gives it a commercial interest in seeing agent deployments continue to grow.

Nvidia has been assembling the pieces for some time, including the earlier open secure AI alliance it announced with partners. Whether the new platform becomes a standard will depend on how many of the named partners ship products built on it, and on whether enterprises trust a hardware vendor to police the agents its own chips run. Nvidia has not published performance figures for Sentry or said when partner products will reach customers.

Disclaimer: AIstify is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, technology, software, and digital innovation sectors. These relationships do not influence AIstify’s editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.

AI & Machine Learning, Cybersecurity & Privacy, News