According to the Financial Times, clinicians are opposing the extension of medical AI tools beyond diagnosis — towards treatment decisions, monitoring or patient triage — citing the weakness of the available performance data. The argument is not a matter of principled distrust: it is a demand for evidence.
We devoted an article to the real state of AI in medicine, distinguishing what is established from what is promised. This stance maps precisely onto that boundary.
Why diagnosis is a favourable case
It is worth understanding why AI progressed first in diagnosis, because that explains where it now stumbles.
Reading a medical image has three properties that make it measurable. The task is bounded: one image, one question. The truth is known after the fact, through a biopsy or the patient's course. And there are vast archives of annotated images available for training and evaluation.
These three conditions make it possible to answer the question that matters: does the system get it wrong more or less often than a specialist? That is why the strongest results concern radiology, dermatology or ophthalmology.
Why the rest is different
Deciding on a treatment has none of these three properties.
The task is not bounded: it depends on the patient's history, their other treatments, their preferences, their social situation. The truth is not known: we will never know what the treatment we did not choose would have produced. And the training data reflect past practice, with its biases, rather than good decisions.
Measuring a system's performance on this task therefore requires what medicine requires of any treatment: clinical trials, with comparable groups, pre-defined criteria and follow-up over time. That is long, costly, and it is precisely what is missing.
One point deserves to be understood, because it separates the two worlds. In technology announcements, a system's performance is measured by the accuracy of its answers. In medicine, it is measured by the state of patients. A tool that produces excellent recommendations on paper but that clinicians follow poorly, or that makes them rely on it to the point of missing what they would otherwise have seen, can worsen outcomes while displaying excellent scores. Only a trial on real patients can establish this.
Caution is not conservatism
It would be easy to read this stance as the reaction of a profession defending its territory. That would be unfair.
Medicine took a century to build the demand for evidence it applies today, and it built it out of disasters: treatments adopted on the strength of their plausibility, which turned out to be useless or harmful once they were finally evaluated properly.
Asking that AI be held to the same standard as any medicine is not hostility: it is the application of a rule that the discipline imposes on itself. We wrote as much about calibration: a system as confident when it is wrong as when it is right is particularly dangerous where an error costs a life.
Key takeaways
The boundary these clinicians draw is not arbitrary: it falls exactly where performance ceases to be measurable by simple means.
That does not mean AI will never have a role beyond diagnosis. It means that role will have to be demonstrated by the methods medicine applies to everything else, and that those methods take years.
This is probably the clearest case where slowness is a method rather than an obstacle. And it is also a reminder that announcements of capability, however impressive, do not replace the question every doctor asks of a new treatment: are patients doing better?