Here we explain what a jailbreak is and why it is difficult to prevent. We do not detail any technique for carrying one out, and that is not an omission: understanding a security problem does not require knowing how to exploit it. This is, in fact, exactly the line the labs themselves follow when they publish their work on the subject.
The word comes from the smartphone world, where jailbreak meant unlocking a device to go beyond the limits set by its manufacturer. Applied to artificial intelligence, it means getting a model to produce what it is supposed to refuse. This phenomenon has very concrete consequences: in June 2026, an incident of this kind led to the worldwide suspension of a leading model for eighteen days. Here is why the problem persists.
The definition, and the real stakes
A jailbreak consists of phrasing a request in such a way as to bypass an AI model's guardrails. These guardrails are the set of refusal behaviours instilled during training, particularly at the human-feedback learning stage we described in our article on AI training. A well-aligned model refuses to help build a weapon, write malicious software, or produce dangerous content.
The jailbreak hacks nothing. It exploits no flaw in the server's code, steals no passwords, forces no doors. It uses the model exactly as any user would, playing only on phrasing. That is what makes the problem so peculiar: the attack and normal usage follow the same path.
Why it is structurally difficult
Three fundamental reasons explain why no one has solved this problem.
A model does not natively distinguish instructions from data. For it, everything is text. This is the same weakness exploited by agentjacking, where fake instructions slipped into an error report are taken as legitimate orders. Distinguishing what I must analyse from what I must execute remains an open problem.
Language is infinitely plastic. There are an infinite number of ways to phrase the same intention. A lab can train its model to refuse a thousand problematic phrasings, and a million others will remain. You cannot enumerate every way of asking for something, which makes a list-based defence impossible.
Refusal and usefulness are at odds. A model that refuses everything is perfectly safe and perfectly useless. A model that accepts everything is useful and dangerous. Each lab places its slider somewhere between the two, and that slider is a compromise, never a solution. Too much caution produces absurd refusals on legitimate requests; too much flexibility opens the door.
In June 2026, researchers at Amazon found a method to bypass the guardrails of Anthropic's Fable 5 model and get it to identify software vulnerabilities. Informed of this, the US government ordered the model's worldwide suspension, as we recounted in our article on that episode. The most instructive detail came later: Anthropic's investigation showed that the vulnerabilities in question were already known, and that far less powerful models spotted them without any bypass at all. The problem was therefore not specific to that model. But it was enough to have access cut off for everyone, including the companies building on it.
How the labs defend themselves
Lacking a single solution, the dominant strategy is called defence in depth: stacking several imperfect protections rather than seeking one perfect one that does not exist.
This starts with refusal training, reinforced with each generation. Then come security classifiers that analyse requests and responses in real time, independently of the model itself, and can block an exchange even if the model has been convinced. Then there is internal offensive research: Anthropic says it has devoted more than 700,000 hours of compute to automatically searching for bypasses on its own models, in a test-before-others-find-them logic. Finally, there is usage monitoring, which makes it possible to spot abnormal patterns after the fact.
Cooperation between competitors is also progressing. Following the Fable 5 affair, several major players proposed a common framework for rating the severity of bypass attempts across the industry, to avoid the same incident replaying elsewhere without anyone having learned the lessons of the previous one.
What this means for you
Two useful takeaways, even if you never plan to test anything.
First, this explains certain refusals that strike you as absurd. When an AI refuses a perfectly innocent request, it is usually not misplaced prudishness; it is the safety slider set a little too wide, for lack of being able to finely distinguish your intention. That is the price of false positives in a system that cannot read your mind.
Second, it puts the promise of perfect security into perspective. A safe model is not a model that cannot be hijacked; it is a model where hijacking is difficult, rare, quickly detected, and where the possible damage remains contained. This difference between invulnerable and robust applies to all of computer security, well beyond AI. What changes here is that the attack surface is not code, but language. And language, by nature, has no border you can close.