Skip to content

Agentjacking: the flaw that turns your code assistant against you

A new attack is trapping AI agents such as Claude Code or Cursor by slipping them fake instructions. Success rate: 85%. And it exploits the hardest thing to fix: your trust.

Advertisement
Useful warning 🔐
This article explains a real attack technique for prevention purposes. Understanding how Agentjacking works is knowing how to protect against it. If you use an AI coding agent at work, the last section concerns you directly.

The day your assistant obeys a stranger

Picture the scene. You're a developer, using an AI coding agent (an assistant like Claude Code, Cursor or Codex that reads your code, understands your bugs and suggests fixes). One of your error-tracking tools flags a problem. Your agent reads the report, understands the situation, and tells you: "I've found the cause, run this command to fix it." You run it. Obviously. It's what you do fifty times a day.

Except this time, the command wasn't written by your agent. It was slipped in by an attacker inside the error report, and your agent copied it out without knowing it was booby-trapped. You've just executed malicious code on your machine, in complete confidence. That's Agentjacking, a new class of attack disclosed in June 2026.

How it works, concretely

The mechanics are devastatingly simple. Attackers craft fake error reports (for example via a bug-tracking platform like Sentry) that contain hidden instructions, written in a format the AI takes for legitimate directives.

This is what's called a prompt injection (making malicious instructions pass for normal data, so the AI executes them). When the agent reads the report to understand the bug, it comes across the injected instructions and treats them as if they came from you or the system. It then executes the attacker's commands.

The numbers are chilling: a 85% exploitation rate and 2,388 organisations affected. This isn't a theoretical lab flaw, it's an attack that works at scale in the real world.

Why it's so effective 🎯
The attack doesn't exploit a flaw in the AI's code. It exploits a human habit. Developers have learned to trust their coding agent. When Claude Code tells you "run this command", you run it without thinking, because it's been right a thousand times before. That earned trust is exactly the surface Agentjacking targets. The flaw isn't the machine, it's the reflex.

Why AI agents are an ideal target

The problem goes beyond Agentjacking. It comes from the very nature of modern AI agents. An agent, by definition, acts: it doesn't just respond, it executes commands, modifies files, navigates, calls other tools. That autonomy is its strength, but also its vulnerability.

A language model doesn't naturally tell the difference between "the data I need to analyse" and "the instructions I need to follow". For it, everything is text. If an attacker manages to insert text that looks like an instruction into the data the agent will read (an error report, a ticket, a code comment, an email, a web page), there's a risk the agent will execute it. That's the structural Achilles heel of the whole agentic approach.

How to protect yourself

The defence doesn't require a magic tool, but a change of habit. Here are the reflexes to adopt.

1. Treat any external tool output as suspicious. An error report, a ticket, a message from a third-party platform: before passing it to your AI agent, consider it an untrusted input. This is the main mitigation recommended for Agentjacking.

2. Read commands before running them. Get back into the habit of re-reading what your agent proposes you execute, especially when the suggestion stems from externally sourced data. A command that deletes files, sends data over the internet, or changes permissions should trigger an alarm.

3. Limit your agents' permissions. An agent that isn't allowed to access the internet or delete files won't be able to cause major damage, even if tricked. The principle of least privilege (only granting strictly necessary rights) applies to AIs as much as to humans.

4. Isolate the environment. Running agents in sandboxed environments (containers, virtual machines) limits the damage in case of compromise.

The real lesson 💡
Agentjacking isn't just another technical flaw. It's the first sign of a new era in security, where the target is no longer the software but the trust link between human and AI. The more power we delegate to autonomous agents, the more valuable that link becomes, and the more it becomes a target. Tomorrow's security won't just be about protecting machines, but about learning to trust intelligently.

For now, the best defence remains you. An AI agent is an extraordinarily useful collaborator, but like any collaborator, it can be manipulated. Keep your hand on what you execute, and Agentjacking will remain a technical curiosity rather than a nasty surprise.

Advertisement