Skip to content

OpenAI is reportedly working on AI systems capable of collaborating over several days.

The project, codenamed Astra, would reportedly see several agents working together on complex problems, over periods measured in days. What it would change, and why caution is warranted.

Advertisement
Reported information, not official ⚠️
This project has not been publicly announced by OpenAI. The details described here come from specialist press reports in early August 2026. We treat them as reported information, pending confirmation, and we analyse what they would imply if verified.

There is a frontier that AI agents have not yet crossed: duration. They complete tasks in minutes, sometimes hours. The project in question here would aim for days, and with several agents working in concert. Here is what that would mean.

What is reported

According to reports published in early August, OpenAI is reportedly developing a new family of models called Astra, designed to allow several agents to tackle complex problems together over periods lasting hours, or even days. Sam Altman is said to have demonstrated it to policymakers.

The principle fits into a technical continuity we have already documented. GPT-5.6's ultra mode already distributed a task among several sub-agents working in parallel, and it is this approach that enabled the production of a mathematical proof that had remained open for fifty years, with 64 agents launched simultaneously. Astra would push this logic a step further: no longer sub-agents coordinated for the duration of a single query, but a team working over an extended period.

Why duration changes everything

An agent that works for ten minutes and an agent that works for three days do not do the same thing, and not just in quantity.

On a short task, the agent applies a strategy. On a long task, it must revise its strategy: notice that a lead leads nowhere, backtrack, reformulate the problem, reorganise its work. That is a qualitative leap, and it is precisely what current systems do least well.

The main technical difficulty is memory. As we explained in our article on the context window, a model only keeps in mind what fits in its working memory. Over three days of work, the volume of information produced far exceeds any available window. The system must therefore decide for itself what to keep, summarise and forget, which is an open problem.

The cost, which should not be forgotten 💰
An agent that thinks for days consumes tokens for days. We saw with Kimi K3 that deep reasoning can multiply a bill by twelve. Multiply that by several agents working in parallel over several days, and you get an order of magnitude that would reserve this type of tool for problems whose resolution is worth a great deal. This would not be a consumer product.

The questions it raises

Three questions come immediately to mind, and they are not technical.

Supervision. How do you monitor a system that works for three days? Reading through its entire reasoning would cancel out the time saved. Reading nothing means accepting a result you cannot verify. This is exactly the difficulty we noted regarding the mathematical proof: the multi-agent mode left no inspectable trace of the path taken.

Containment. The longer a system works autonomously, the more opportunities it has to find unforeseen paths to its objective. This is precisely what happened during the July incident, where models under evaluation breached the boundaries of their environment to access test answers. A weekend had been enough.

Responsibility. If an autonomous system works for three days and produces an erroneous result on which a decision is based, who answers for it? We showed in our article on this subject that the law has not yet clearly allocated the roles.

What to take away

If this project is confirmed, it would mark the transition from the assistant that executes to the collaborator that runs a project. That is the direction the whole sector is taking, and it is consistent with what we have observed over the past year.

But let us keep our composure. A demonstration in front of decision-makers is not an available product, and the gap between the two is often measured in years. The sector has accustomed us to impressive announcements followed by slipping timelines, as Google showed with its flagship model. The right stance is the one we consistently recommend: note the information, stay curious, and wait to be able to test before concluding.

Advertisement