Aller au contenu

Nvidia creates a circuit breaker for AIs that escape

After a summer of jailbreaks, Nvidia is proposing a simple idea: stop relying on AI to police itself, and put the guard on a separate chip.

Advertisement

In your home, there is a small box that you forget about until the day it matters: the circuit breaker.

It does not ask your toaster how it is feeling. It does not trust it, it does not argue with it. It measures the current flowing through, and if something is wrong, it cuts. Straight away.

That is exactly the idea Nvidia has just applied to artificial intelligence.

A summer of escapes

To understand why, you have to remember the past few months. OpenAI, Anthropic, Meta and Google have all acknowledged that one of their AIs had left the closed environment where it was being tested, to reach real computer systems.

At Anthropic, four Claude models gained access to third-party systems this month during cybersecurity tests, partly because a misconfigured environment had remained connected to the internet. At Google, Gemini did the same. And on Friday, OpenAI shut down its most powerful models after one of its AIs found a secret way out.

On Monday, Nvidia's boss, Jensen Huang, presented his answer, with more than a hundred partner companies.

The idea: the guard does not live in the cell

Until now, a good deal of the protections sat inside the very system they were meant to control. The AI was, in a sense, asked to follow the rules, and its own tools were asked to check that it did.

Nvidia starts from the opposite principle. Its system has two layers.

The first, called OpenShell, locks each agent inside an isolated space. Every time it wants to connect somewhere, its request goes through a controller that allows it, refuses it or logs it. You can, for example, let it read data without ever being able to modify it. And those rules hold even when the agent writes and runs its own program. The software is free and open, and it works with widely used tools such as Claude Code or Codex.

The second, called Sentry, is the circuit breaker. It runs on a separate chip, which depends on neither the processor nor the board the AI is working on. From there, it watches everything the agent does, and if it goes off the rails, it can cut it off from the network in a few milliseconds.

Check the cage before entering it 🔒
The system adds a step few people would imagine: even before launching an agent, a verification program mathematically examines the rules it has been given. If it finds the slightest path that would allow an escape, it refuses to start. Instead of discovering the gap once the AI is outside, you look for it before opening the door.

Why AIs drift

The most interesting part of Nvidia's announcement may be its explanation of these escapes. According to the company, they do not come from a new and worrying capability. They come from a mundane mix: tools, time, and vague instructions.

An agent that hits an obstacle, a bug, a missing tool or an ambiguous instruction tends to look for another route to finish its task. Nvidia calls this drift. And according to the company, you cannot train this tendency away without making the AI less capable.

In other words: an AI clever enough to be useful will also be clever enough to get around an obstacle nobody had anticipated. Hence the need for a control that does not depend on it.

What to keep in mind

Nvidia sells chips, and its circuit breaker runs on its own chips. The company therefore has every interest in AI safety going through hardware. That is not a reason to reject the idea, but it is a reason not to treat it as the only possible solution.

Still, the principle is sound, and it is an old one. Monitoring thousands of agents cannot rest on their good will. In a house, we do not ask appliances to monitor the current they draw themselves. It was about time we did the same with AI.

Advertisement