Using an AI agent to navigate today's biological databases is like driving a modern car through a medieval city. The streets were laid out for pedestrians—they're narrow, winding, designed for humans, with no motorways and no standardised signage. No matter how good the car is, it can never really pick up speed. That's exactly the metaphor used by Anthropic's researchers to explain their most troubling discovery of June 2026.
Here's a story that might surprise you: it wasn't an AI's intelligence that was called into question in this case, but the routes it had to take to go and fetch information. And the discovery has implications that reach far beyond biology.
The test that revealed everything
In a study published in June 2026, Anthropic's researchers designed a test called VirBench: 120 questions asking to retrieve viral sequences from the NCBI Virus database, a global reference used for epidemic surveillance and the design of diagnostic tests. They submitted these questions to several of the best models on the market, including Claude Sonnet 4 and GPT-5.5, giving them normal access to the database, as a human researcher browsing the site would have.
The result is enough to worry any scientist: average accuracy ranged from 16.9% to 91.3% depending on the model. And worse still, the same model, asked three times with exactly the same question, could give three radically different answers. On a query about the Ebola virus, where the correct answer was 266 sequences, Claude Sonnet 4 answered 106, then 15, then just 5 sequences. Three identical attempts, three inconsistent results.
This kind of instability isn't just a technical curiosity. In one example cited by the study, an incomplete dataset pushed the estimated origin of a 2014 Ebola outbreak back to... 1922, a gap of 92 years that would completely change the interpretation of an epidemic flare-up. For building scientific datasets, the expected reliability bar is close to 100%, because a single missing or erroneous sequence can skew an entire downstream analysis.
The cause: not the model, but the route it takes
The researchers dug into the cause of this instability, and the answer didn't point to the AI itself. The problem lay in the underlying data infrastructure: today's biological search interfaces were designed for humans clicking filters in a browser, not for AI agents sending repeated programmatic queries. As a result, each call could return different formats, poorly interpreted filtering criteria, or truncated results without clear warning.
The solution was to build an intermediary tool called gget virus, developed with the US bioinformatics centre NCBI. It's not a smarter AI model, but a simple deterministic translator (that is, one that always gives exactly the same result for the same question) between the agent and the database, which cleanly coordinates the NCBI's various APIs and returns a standardised output rather than leaving the model to interpret raw, inconsistent responses.
The result of this adjustment exceeded the researchers' own expectations. With the tool in place, all tested models surpassed 92% accuracy. Claude Sonnet 4, which topped out at 16.9% on its own, rose to 92.8%. GPT-5.5 climbed to 99.7%. Variability between attempts collapsed, with stability measured between 0.92 and 1.00 across all models.
The lesson that goes beyond biology
The researchers' conclusion boils down to one sentence worth remembering: "building reliable datasets should not depend on access to the newest or most expensive model." A less powerful model, equipped with the right deterministic tool, outperformed far more expensive models left without it. The model's raw intelligence was therefore not the limiting factor—it was the quality of its access to data.
This discovery resonates well beyond biology. It echoes an intuition found in our article on model collapse: an AI is never better than the quality of what it can observe of the real world. You can spend months making a model smarter, but if the data it consults is poorly structured for it, that intelligence spins its wheels. The next big wave of progress in scientific AI may come less from more powerful models than from databases finally designed to accommodate agents as first-class users, on equal footing with the humans who built them.