A field report published by OpenAI with academic partners shows that coding agents can take over and modernise abandoned scientific software, with speed gains of up to a factor of 60. The same study delivers a direct warning: these systems are eloquent, convincing, and sometimes categorically wrong. Both findings matter equally.
There is a problem in scientific research that gets little attention: a considerable amount of essential software was written years ago by PhD students who have since moved on, runs on obsolete technologies, and is no longer maintained by anyone. This report describes what happens when you set AI agents loose on that task.
The problem it solves
A significant share of modern science depends on code. Simulations, processing of observational data, statistical analysis: these tools are often developed by researchers who are not software engineers, to meet a specific need, then abandoned when their author changes labs.
The result is a substantial body of software that works poorly, runs slowly, and that no one understands anymore. Rewriting it interests no funder, because it is not a discovery. It is an unrewarding, essential task that is systematically postponed.
This is exactly the kind of work where a coding agent excels: well-scoped, technically demanding, requiring no scientific creativity. The reported gains—up to sixty times faster on certain workloads—concretely change what a lab can afford to compute.
The warning, and why it is central
The phrase used by the authors deserves to be remembered: these systems are eloquent, convincing, and confident in error. This is not a polite caveat; it is the heart of the problem.
We explained the mechanism in our article on calibration: an AI produces a wrong answer with exactly the same confidence as a right one. Applied to scientific code, this takes on a particular dimension.
Code that crashes is a visible problem. Code that runs perfectly and produces a slightly wrong result is an invisible problem, and it can propagate into publications, datasets, and the work built on top of them. A sixty-fold speed-up is useless if it produces erroneous results sixty times faster.
We recounted in an article on VirBench how a model asked the same biology question three times gave three different answers, none of which was correct. The solution had not been to use a smarter model, but to give it a deterministic tool. The lesson repeats itself here: in a scientific context, reproducibility matters more than performance.
What this implies for the scientific method
This tension raises a question that goes beyond computing. Science rests on verifiability: a result only exists if someone else can retrace the path. Yet code rewritten by an agent—faster and cleaner, but with no human having reviewed every line—introduces a blind spot at the heart of the apparatus.
The risk is not theoretical. A researcher who obtains a result consistent with their expectations, produced by a tool that looks competent, will have little reason to dig deeper. This is the bias we described in our verification method: we scrutinise far less what suits us.
The practices that address this difficulty exist, and they are not new: keep the old code as a reference, compare outputs on known cases, require results to be reproduced by a third party. What changes is that these precautions move from being good practice to being an absolute necessity.
What to take away
This report is interesting precisely because it does not pick a side. It documents a real and considerable benefit, and it names the risk without downplaying it. That is rare in a sector where corporate publications rarely highlight the limitations of their own products.
The practical conclusion extends well beyond laboratories. Every time an AI saves you a significant factor on a task, ask yourself what verifies the result. If the answer is nothing, because it looks right, you have not saved time: you have moved a risk to a place where no one is watching it.