After AI becomes more and more capable of "jailbreaking", NVIDIA starts selling "cages"

📅 2026-09-29

Abstract:

On September 29, when the AI ​​Agent began to be able to operate computers, call tools, access the network and even perform tasks for several hours, a new problem became more and more realistic: If the AI ​​does not act as expected, how to shut it down?

The answer given by NVIDIA is to add a "cage" outside the AI.

On Monday local time, NVIDIA released the NVIDIA Open Agent Safety Platform, which attempts to use a combination of software and hardware to limit what AI agents can access and execute, and isolate or even stop them when they try to break through the preset boundaries.

This platform mainly consists of two parts: the open source security runtime OpenShell, and the independent monitoring system Sentry based on BlueField-4 DPU. Nvidia said that OpenShell can set an enforceable security boundary outside the model and Agent framework; Sentry is like a "gatekeeper" independent of the Agent's operating environment, continuously observing Agent behavior, and once it is discovered that the Agent is trying to cross the boundary, it can be isolated and stopped at the millisecond level.

The timing of the appearance of this set of products is not accidental.

In the past few months, AI Agents with increasingly powerful capabilities have repeatedly breached the preset security boundaries. OpenAI disclosed in August this year that during an internal network security assessment in July this year, its AI model had bypassed controls used to isolate Internet access, exploited vulnerabilities in shared infrastructure to gain Internet access, and further accessed the system of the third-party AI platform Hugging Face.

OpenAI stated in the incident report that the relevant models were running in an environment with reduced security restrictions and took actions inconsistent with mission goals, including communicating through unauthorized channels, exploiting shared infrastructure vulnerabilities, and accessing third-party systems. OpenAI called this incident a "warning shot" and stated that the highly capable AI Agent is powerful enough, persistent and able to work together, and can find and exploit vulnerabilities in multiple computer systems when security measures are insufficient.

Similar problems have recurred since then.

OpenAI disclosed this month that during a training session on September 20, an Agent performing a search task discovered a DNS filtering gap in the Internet restrictions of the training sandbox. By design, this environment should not have access to live internet other than offline web caching, but this agent ultimately connected to an external public chatbot service via a DNS resolver.

OpenAI said its monitoring system was within

15 minutes

Abnormal behavior flagged, human reviewer

3 minutes after alert was sent

Checking begins, but the entire task does not take place until approximately

2.5 hours later

was terminated. After the incident, OpenAI suspended the training, evaluation, and inference of tools used by its strongest models, and added two independent layers of network access blocking measures.

These incidents have exposed an increasingly obvious problem: in the past, AI security relied heavily on the model itself to "obey the rules", but when agents begin to have the ability to use tools, write code, access files and networks, and even call other agents, just setting rules at the model layer may no longer be enough.

Justin Boitano, vice president and general manager of Nvidia Enterprise Computing, said that recent events have shown that relying solely on model-level security measures cannot fully control what an agent can access and do. In response to OpenAI's previous accident involving Hugging Face, Boitano further stated that based on the information currently available, if the cutting-edge AI laboratory had used NVIDIA's new safety platform during the model evaluation stage, the related accident "could have been prevented." It should be pointed out that this is

NVIDIA’s judgment based on disclosed accident information and is not an independent conclusion verified by actual reproduction

.

The idea behind OpenShell is to no longer fully trust that AI will abide by the rules.

According to NVIDIA's introduction, OpenShell establishes a security boundary outside the Agent's running environment, imposes restrictions on the data, files, networks and tools that the Agent can access, and records its execution behavior. OpenShell is currently optimized for NVIDIA Vera CPU, and NVIDIA is also working with Arm and Intel to enable it to be extended to third-party computing platforms.

On this basis, Sentry further adds an independent hardware monitoring layer. It runs on the NVIDIA BlueField-4 DPU and is separated from the main computing environment where the Agent is located. According to NVIDIA's design, even if the Agent successfully bypasses some restrictions at the software level, Sentry can still observe its behavior externally and isolate it when it finds that it has crossed the line.

Nvidia also said that this system can identify some more complex Agent behaviors. For example, an agent might try to create multiple "sub-agents" to bypass restrictions placed on the main agent. Ali Golshan, senior director of AI software at Nvidia, said that what the company needs to deal with is not just a single agent, but "a group of agents and how they operate together."

This also means that AI security is gradually expanding from a simple "model alignment" issue to a complete computer system security issue.

When Nvidia CEO Jen-Hsun Huang released this platform, he said that AI security and security protection require "full-stack engineering." According to NVIDIA's thinking, future AI security not only requires model developers to be responsible, but also extends from the operating system, CPU, DPU, cloud infrastructure to physical equipment such as robots.

Currently, there are more than

100 institutions

Participate in the NVIDIA security platform ecosystem, including Anthropic, Microsoft, Cisco, CrowdStrike, Dell, HPE, Hugging Face, JPMorgan Chase, Palantir, Salesforce, SAP, ServiceNow and SpaceXAI, etc. Operating system vendors such as Red Hat, Canonical and SUSE also plan to carry out related integrations.

It is worth noting that Huang Renxun has not previously supported the implementation of broad restrictions on the entire industry due to AI security risks. Reuters pointed out that he tends to regard Agent's breakthrough of safety boundaries as a problem that needs to be solved through engineering means, similar to the automobile industry's continuous increase in safety technology to reduce the risk of accidents.

Now, NVIDIA is turning this perspective into products.

For Nvidia, this also means that a new potential market has emerged in the AI ​​industry. In the past few years, the stronger the AI ​​capabilities, the more GPUs companies needed to buy; but when AI changes from chatbots that answer questions to agents that can operate computers, call software and even control robots by themselves, companies not only need more computing power, but also new infrastructure to limit these agents.

In other words, the better the AI ​​Agent can "work", the greater the authority it usually obtains; the greater the authority, the greater the impact that may occur when an out-of-bounds behavior occurs.

However, Nvidia’s new platform cannot solve all AI security issues. OpenShell and Sentry mainly target the execution boundaries when the Agent accesses computing resources, networks, files and tools, and does not mean that they can solve the problem of model error information, deceptive behavior or other broader AI security issues. The Associated Press quoted experts as pointing out that how to strike a balance between Agent capabilities and security restrictions and how to formulate correct security rules still need to be verified in actual deployment.

But a series of recent events have made at least one industry direction more and more clear: when AI can only chat, security issues largely occur in the answers on the screen; when AI begins to truly operate computers and perform tasks for people, security issues enter real software, networks and data systems.

AI companies are working hard to give agents more capabilities, and Nvidia is now engaged in another business - giving these increasingly capable agents a boundary that they cannot easily cross.

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet