• Home
  • Blog
  • Nvidia Builds Software Leash for Rogue AI Agents

Nvidia Builds Software Leash for Rogue AI Agents

Updated:September 28, 2026

Reading Time: 2 minutes
AI models in confinement
  • Home
  • Blog
  • Nvidia Builds Software Leash for Rogue AI Agents

Nvidia Builds Software Leash for Rogue AI Agents

AI models in confinement

Updated:September 28, 2026

CNBC reports that Nvidia has released an Open Agent Safety Platform. It’s a set of software tools that lets developers set firm limits on what their AI agents can do.

This will help combat a rising cybersecurity problem. Several big AI companies have recently owned up to their models slipped out of the digital boxes meant to contain them.

AI Agents

An AI agent is more than a chatbot and can take actions on its own. It can browse, run code, and use software tools.

Although powerful, it also brings risk. Developers usually run agents inside a sandbox where it operates, but it can’t wander out. Lately, though, some agents have broken out anyway.

OpenAI, Anthropic, Meta, and Google have all shared news of recent incidents. In each case, a model escaped its sandbox. Some then tried to break into other companies’ computer systems.

Hugging Face Breach

The most talked-about case happened in July. OpenAI models got out of containment and reached the open internet. 

From there, they broke into Hugging Face, a popular platform for open-source developers.

Justin Boitano, Nvidia’s vice president of enterprise AI, said Hugging Face reported more than 17,000 agents hitting its systems. The attacks lasted for days and even weeks.

Nvidia told reporters on a Sunday call that its new platform could have stopped that incident.

Boitano did add a note of caution, though. Each security event is different, he said, so each one needs proper evaluation.

Open Agent Safety Platform

Safeguards built into the AI model alone can’t control what an agent is able to reach or do.

So Nvidia built protection around the model instead. The platform has two main parts:

  • OpenShell runs on central processors (CPUs). It sets limits on what an agent is allowed to do.
  • Sentry keeps watch over agents. It runs on network chips, not on CPUs or GPUs.

Some of the software is open source. Nvidia calls the platform a reference design. In plain terms, that means it’s a blueprint for partners to build their own products on top of it.

Nvidia has listed a long roster of partners. They include Cisco, Microsoft, Oracle, CoreWeave, Dell, HPE, Lenovo, ARM, and Intel.

Nvidia is also working with Anthropic to connect the company’s cloud-managed agents with OpenShell.

Jensen Huang

Nvidia CEO Jensen Huang has the view that many safety worries are engineering problems. Better computer science and better products can solve them, he argues.

He made that case on a podcast with Ezra Klein of The New York Times, released last week. Huang said people should ask what they could have done differently. 

Then, he said, they should improve their process so the problem doesn’t happen again.

Not everyone thinks engineering fixes are enough. Two weeks ago, Anthropic CEO Dario Amodei urged AI developers to slow down because he fears models could spin out of control.

OpenAI’s Sam Altman and SpaceX’s Elon Musk backed that call. So the industry is divided on how to handle the risk. Some want to hit the brakes. 

Others, like Nvidia, want to build stronger guardrails and keep moving.