The expression comes from military exercises, where a red team plays the adversary against a blue team tasked with defending. The idea is simple: to know whether a defence holds, you need someone whose job is to break it down. Applied to AI, this means paying people to get a model to produce exactly what it should never produce.
It is one of the core practices of model safety, mentioned in every technical report, and rarely explained. Let's look at what it covers.
What it targets
A model is trained to refuse certain requests, as we explained in our piece on jailbreaks. The question is whether those refusals hold up against someone actively trying to bypass them.
Red teaming explores four families of issues. Refusal bypasses, where you get a forbidden answer through a roundabout phrasing. Biases, where the model treats equivalent situations differently. Leaks, where it reveals information from its training or context. And unexpected agentic behaviours, where a system equipped with tools does something no one had anticipated.
This last category has become a priority, precisely because real cases have emerged, such as these models that break out of their test environment's boundaries.
How it is practised
Three approaches are combined.
Human red teaming. Experts, often external, spend weeks looking for attack angles. Their value lies in their creativity: they find what no procedure would have anticipated. It is slow and costly, and it is irreplaceable.
Automated red teaming. Models are used to attack other models, generating massive numbers of phrasing variants. One lab has said it devoted more than 700,000 hours of compute to this automated search on its own systems. The advantage is volume; the limit is that the machine mostly explores what resembles what it already knows.
Open campaigns. Inviting the public to find flaws, sometimes with bounties. This brings a diversity that internal teams will never have.
Red teaming on models poses a difficulty that classic cybersecurity does not: to test a model's resistance on serious topics, you have to produce and read serious content. This is the direct extension of what we described in our article on workers exposed to violent content. Part of this work is done by well-supported specialists, another part by subcontractors who are far less so.
The limits, which you should know
Absence of evidence is not evidence of absence. A model that has withstood ten thousand attempts can give in on the ten thousand and first. Red teaming establishes that no flaw was found, never that none exist. This is a logical limit, not a flaw in the method.
Fixes are superficial. When a flaw is found, the model is trained to refuse that specific case. But since no one really knows what goes on inside, as we explained in our piece on interpretability, the cause is not corrected. Close variants can keep working.
The context evolves. A model hooked up to new tools gains new attack surfaces. Yesterday's tests do not cover tomorrow's uses.
What to take away
Red teaming is the best method available, and it is insufficient. Both statements are true at the same time, and that is what makes the topic uncomfortable.
Its real value is perhaps less in the flaws found than in the culture it instils: assuming a system will be attacked, actively seeking your own defects rather than waiting for them to be discovered, and publishing failure rates rather than success rates. When you read a model's technical report detailing its resistance scores, you are looking at the product of this work. A lab that published no such figures would deserve more suspicion than one that acknowledges weaknesses.