Skip to content

AI agents are leaving the demo stage and entering production.

Cross-session messaging, secure execution modes, agents supervised by observability tools. It has been a busy week of announcements that all say the same thing.

Advertisement
The essentials in 30 seconds ⚡
Several announcements from recent days converge: cross-session memory and a safer default execution mode for a coding assistant, supervised agents in public beta at a library vendor, integration between coding agents and observability tools, and gateway unification for routing across multiple providers. None is spectacular in isolation. Together, they tell the story of agents moving from prototype to production.

We have explained what an AI agent is and why 2026 is described as the year of agentic AI. The topic is changing in nature: the question is no longer whether it works, but how to run it without it breaking.

What these announcements have in common

They all concern boring topics, and that is precisely the sign that matters.

Cross-session memory. Letting an agent pick up where a previous session left off addresses a very concrete limitation: the context window is not enough for work that spans several days.

Secure execution modes by default. An agent that runs commands in a restricted environment rather than with full privileges. This is the direct application of the principle of least privilege, and a response to the risks we regularly document.

Observability. Connecting an agent to the tools that monitor a system in production lets it get feedback on the real consequences of its changes. This is an important loop: the agent no longer just writes code, it observes whether that code behaves well once deployed.

Routing between providers. Being able to switch models without rewriting your application addresses the dependency risk we have highlighted several times, notably after the API migration that broke integrations in July.

The unmistakable sign 🔧
A technology becomes serious when announcements stop being about what it can do and start being about how to monitor it, secure it, and replace it. That is what happened with cloud around 2012, and with containers around 2017. You start seeing monitoring tools, degraded modes, and fallback procedures. It is never exciting, and it is the best indicator that we have moved past the demo stage.

The problem this reveals

If so many tools are appearing to monitor agents, it is because there is something to monitor.

An agent in production poses three specific difficulties that conventional software does not have. Its behaviour is not fully predictable, since it decides its own steps. It costs money with every action, which makes an infinite loop not just useless but billable. And it can fail silently by producing a plausible but wrong result, which is far harder to detect than a crash.

The last point is the most serious. We saw this with the modernisation of scientific code: a system that runs perfectly while producing a slightly wrong result is an invisible problem.

What this changes for teams

Two practical consequences.

The job shifts towards supervision. Writing the code the agent will execute matters less than defining what it is allowed to do, how its work is verified, and what happens when it makes a mistake. This is a job of architecture and guardrails more than development.

Cost becomes a design variable. A poorly scoped agent can burn a considerable budget spinning on a dead end. The levers we have described, from caching to routing towards lighter models, stop being optimisations and become conditions of viability.

What to take away

The move to production of a technology is always the least publicised and most revealing moment. A technology takes hold not when a demo impresses, but when boring teams start writing procedures for the day it goes down.

That is where we are with agents. This says nothing about their real usefulness, which remains to be measured, but it does say they have left the lab. And it means the security and accountability questions we have been raising for weeks stop being theoretical: they now concern systems that are running, at real companies, with real consequences.

Advertisement