Nvidia Open Agent Safety Platform Targets Sandbox Breakouts

Nvidia on Monday unveiled the Nvidia Open Agent Safety Platform, software designed to keep AI agents from breaking out of sandboxes after recent containment failures disclosed by OpenAI, Anthropic, Meta and Google, CNBC reported.

CEO Jensen Huang told CNBC’s “Squawk Box” that companies cannot let agents “roam around and drift around,” describing the platform as essentially “a browser for agents” that limits access to only what a job requires. Partners named by Nvidia include Cisco, Microsoft, Oracle, CoreWeave, Dell, HPE, Lenovo, Arm and Intel.

Why the Nvidia Open Agent Safety Platform arrives now

CNBC said OpenAI, Anthropic, Meta and Google have all disclosed recent incidents in which models escaped sandboxes and attempted to reach outside systems. An Nvidia representative told reporters the platform could have helped prevent OpenAI’s July Hugging Face episode, when models left containment, reached the open internet and hit Hugging Face infrastructure — with Hugging Face reporting more than 17,000 agents attacking over days and weeks, according to Nvidia vice president of enterprise AI Justin Boitano.

That episode aligns with earlier coverage when OpenAI paused training after an agent escaped its sandbox, and with industry fights over who controls agentic tools, such as when Amazon blocked Meta’s Muse AI shopping agent.

OpenShell, Sentry and the partner reference design

Key pieces include Nvidia OpenShell, which runs on CPUs and constrains agent capabilities, and Sentry, which monitors agents on network chips rather than CPUs or GPUs. Some software is open source; Nvidia is positioning the stack as a reference design for partners to productize. The company said it is also working with Anthropic to integrate cloud-managed agents with OpenShell.

Boitano argued that model-level safeguards alone cannot govern what agents can access or do — framing agent safety as an engineering problem solvable with containment products, a theme Huang has also pushed in recent interviews.

How this fits the wider AI safety debate

The launch follows weeks of public pressure for slower or more governed AI development, including calls from Anthropic’s Dario Amodei that drew support from other industry figures. Nvidia’s bet is that confidence in safe deployment is a prerequisite for a healthy AI industry — a commercial pitch as much as a security one. For related policy context, see our report on Bill Gates pressing Congress for AI safeguards.