Skip to content

What is an "AI agent"? The word everyone uses without really explaining it

Sonnet 5, GPT-5.6, LongCat-2.0: all of them boast of being "agentic". But what really sets an agent apart from a plain chatbot?

Advertisement
The starting analogy 🤖
A classic chatbot is an employee who answers your questions over the phone, but can do nothing other than talk. An AI agent is that same employee, but you've given them the keys to the office, access to the printer, the accounting software, and permission to go and find the information they need themselves before answering you. The difference isn't in what they can say, but in what they can do.

If you've been following AI news recently, one word appears in absolutely every launch: agentic. Claude Sonnet 5 is presented as "the most agentic Sonnet ever released". GPT-5.6 is betting on its reasoning modes for agentic work. LongCat-2.0 is built for agentic use. The word has become so ubiquitous that it ends up meaning nothing precise anymore. Let's put it back on clear rails.

The simple definition

An AI agent is a system that doesn't just answer a question in a single exchange, but can plan multiple steps, use external tools (a browser, a terminal, a database), observe the results of its actions, and adjust its strategy accordingly, until it completes a full task, often without human intervention at each step.

The difference from a classic chatbot comes down to one image: a chatbot responds, an agent acts. Ask a chatbot to help you book a train ticket, and it will explain how to do it. Ask the same of a well-equipped agent, and it will open the site itself, fill in the form, check the schedules, and only come back to you if it needs a decision (like choosing between two trains).

The mechanism: the agentic loop

Technically, an agent works on a repeating cycle: observe, plan, act, verify. It looks at the current situation (the code it needs to fix, the web page it needs to read), decides the next step, executes a concrete action via a tool, then observes the result to decide what comes next. This cycle can run for a few seconds or for several hours, depending on the complexity of the task.

That's exactly what a tester of Claude Sonnet 5 described in our article on its launch: asked to investigate a bug, the model, without being explicitly told to, wrote a test to reproduce the problem, implemented a fix, then checked that the bug didn't come back. Three steps chained together, decided by the model itself, from a single initial request.

What unlocked agents in 2026 🔓
Two technical advances explain the explosion of agents this year. First, tool use: models can now reliably call external functions, rather than just generating text. Second, much larger context windows, which allow an agent to keep track of everything that has happened during a long task, without losing the thread after a few minutes. Without these two ingredients, agentic AI would remain a lab concept.

Why it's not without risk

Giving an AI the ability to act, not just to talk, fundamentally changes the nature of the risk. An AI that gets an answer wrong produces a false sentence. An agentic AI that gets something wrong can execute a destructive command, send an email to the wrong person, or spend real money. That's exactly the territory exploited by agentjacking, the attack that traps coding agents by slipping fake instructions into data they believe is legitimate.

It's also for this reason that labs apply different guardrails depending on the level of autonomy granted. The more power to act an agent has, the stricter the security mechanisms around it must be, a principle seen in how Anthropic built different security levels between Sonnet 5, Opus 4.8 and Mythos.

What to remember

Behind the buzzword, there's a real transformation: AI has moved from the status of an assistant that answers to that of a collaborator that executes. It's this shift that explains why 2026 is unanimously described as "the year of agentic AI". Next time a lab touts its "agentic" model, the real question to ask isn't "does it answer well?", but "how far can it act alone, and who checks what it's doing in the meantime?".

Advertisement