A research paper shows that encrypted reasoning blocks emitted by several major providers are interchangeable between sessions, between users, and between models within the same ecosystem. The upshot: an attacker can inject the encrypted reasoning of a powerful model into a weaker model from the same family and get it to produce the content in plaintext, without having to bypass the powerful model's protections. The authors demonstrate this across three major providers.
We've covered several security vulnerabilities this year. This one is different in nature, and particularly elegant in its mechanism.
The context: why reasoning is encrypted
Recent models produce an internal draft before answering, as we explained in our article on reasoning models. This draft often contains elements providers prefer not to expose: details about the model's method, sometimes fragments of sensitive information.
Several providers have therefore chosen to transmit these blocks in encrypted form. The user receives an opaque block that the system can reuse later in the conversation without the content being readable.
The flaw
The problem discovered is that these blocks are not tied to the session, nor to the user, nor to the specific model that produced them. They are valid across the provider's entire ecosystem.
The attack follows directly: you grab an encrypted block produced by a powerful model, feed it as context to a weaker model from the same family, and ask the latter to reproduce it. The small model, which lacks the same guardrails, complies.
The trick is that you never attack the protected model. You use a less guarded member of its own family as a translator. It's a variant of the principle we described regarding jailbreaks, but one that sidesteps the question of bypassing entirely.
The authors report having decoded more than three hundred thousand reasoning blocks collected from public repositories, and having extracted from them several hundred pieces of personal data as well as access credentials. This last point is the most concerning: developers have unknowingly published blocks containing secrets, thinking they were unreadable. Encryption created a false sense of safety that led to publishing what would never have been published in plaintext.
The second attack, more insidious
The authors also describe an invisible injection path. Since these blocks are accepted as legitimate context, you can place instructions in them that persist in an agentic execution without being visible.
It's a variant of agentjacking, with an added difficulty: the vector is an object the system considers its own, encrypted, and therefore a priori trustworthy.
The general lesson
This discovery illustrates a classic security principle: encryption only protects if the key is properly tied to a perimeter. Encrypted content that any member of the ecosystem can decrypt is not protected; it's simply unreadable to those without the right door.
It also ties into what we wrote about model opacity: making reasoning inaccessible to the user for reasons of industrial property creates an object whose content no one can verify, including the one carrying it.
What to do
For a developer, two immediate precautions. Never publish encrypted reasoning blocks, in a repository, a journal, or a bug report: treat them as sensitive data, not as opaque tokens. And do not consider a received encrypted block as a reliable source in an agentic chain.
For everyone, this affair recalls something useful: the strongest protections rarely fall by force. They fall because someone finds a side door no one thought to close.