Skip to content

One link was enough to leak your emails: the Copilot flaw that makes agentjacking a reality.

An undocumented parameter, automatic URL retrieval, poisoned persistent memory. Three innocuous building blocks, a silent exfiltration.

Advertisement
The essentials in 30 seconds ⚡
Microsoft published a patch on 18 August for a vulnerability identified by security researchers. The flaw chained together three elements: an undocumented URL parameter, the assistant's built-in ability to fetch page content, and a persistent memory that could be poisoned. The result: a single malicious link could automatically trigger instructions and exfiltrate data from connected services. The patch is available.

We have been describing agentjacking for months as a theoretical risk documented in the lab. Here is a concrete case, in a product used by millions of people.

The mechanism, step by step

What makes this flaw instructive is that none of its building blocks is abnormal in isolation.

An undocumented parameter. A feature present in the product but absent from public documentation. This is not rare in complex software, and it is a classic blind spot.

Automatic content retrieval. The assistant can go read a web page to answer. That is a useful feature, and it is the entry point: the retrieved content is not controlled.

Persistent memory. The assistant keeps information from one session to the next, as we explained in our article on what your AI knows about you. If you manage to write to that memory, the effect does not disappear at the end of the conversation: it persists.

The chain produces a one-click attack. The link is opened, the content is read, hidden instructions are interpreted as legitimate, and they execute with the user's rights on their connected services.

Why this is structurally hard to prevent 🔓
We have said it before: a model does not natively distinguish the data it should analyse from the instructions it should follow. It is all text. When an assistant reads a page, it has no reliable mechanism to decide that a sentence found on that page is not an order from its user. This is the same problem as the one laid out in our article on jailbreaks, and there is no general solution to it to date.

The aggravating factor: memory

This is what sets this flaw apart from previous attacks. A classic injection acts for the duration of a session. An injection that reaches persistent memory settles in.

Concretely, a malicious instruction can keep acting days later, in unrelated conversations, with nothing recalling the link originally opened. The victim cannot link cause to effect.

This raises a broader design question. Memory genuinely improves the experience, and it mechanically widens the attack surface over time. The two go together.

What to take away from this

Apply the patches. The advice sounds mundane; it is the most effective. We documented with a campaign targeting 460 systems that attacks massively exploit already-patched vulnerabilities. The delay in applying them is the main window of risk.

Look at what you have connected. An assistant linked to your email and your documents has a considerable scope. The question to ask is not what the tool brings you, but what it could reach if it were hijacked.

Inspect your memory. Interfaces generally let you review what is retained. That is good hygiene, and this episode gives it an extra reason.

What to remember

This affair moves agentjacking from the status of documented risk to that of a patched incident in a consumer product. It is an expected step, and there will be others.

It also confirms what we wrote about agents entering production: the more these systems gain in action capability, connectors and memory, the more their attack surface expands. Every feature that makes them useful also makes them exploitable. There is no version of these tools that is both powerful and without a surface.

Advertisement