This article provides a general overview and does not constitute medical advice. It does not address any individual situation. For any personal health questions, the opinion of a professional remains essential, and no consumer AI tool is a substitute for it.
Few fields produce as many spectacular announcements and as much slow deployment. Let's sort out what actually works in services from what remains at the publication stage.
What already works and is being rolled out
Imaging. This is the most mature field, by far. Detecting an anomaly on an X-ray, mammogram or CT scan is a pattern recognition task, exactly what these systems do best. Several tools are approved and used in routine practice, generally in second reading: the radiologist analyses, the tool flags what they might have missed. This is the most solid setup, because it adds safety without removing responsibility.
Administrative work. This is the most underestimated benefit and probably the most immediate. A considerable share of medical time goes into reports, letters, procedure coding. Automatic transcription of a consultation and assisted report writing give clinical time back. No diagnostic prowess here, but a concrete effect on caregivers' availability.
Triage and prioritisation. Spotting, in a queue of exams, those showing signs of urgency allows the reading order to be rearranged. Here again, the tool does not decide, it draws attention.
Research. Analysis of massive datasets, identification of drug candidates, help with trial design. This is probably where the long-term impact will be greatest, and it is also the least visible to the public.
What does not work as advertised
Headlines like "AI beats doctors at such-and-such exam" deserve to be read with caution, for three reasons.
An exam is not a consultation. Answering a written clinical case, with all the elements provided and an existing correct answer, has little in common with a real consultation where the patient describes their symptoms poorly, forgets details, and where the bulk of the work consists of asking the right questions.
Performance drops outside the lab. A system trained on one hospital's data often works less well in another, with different equipment, a different population, different protocols. This gap is the main obstacle to deployment, and it is systematically underestimated.
The absence of an uncertainty signal. This is the most critical point in medicine. As we explained in our article on calibration, these systems assert with the same confidence whether they are right or wrong. In a context where knowing that you don't know is a fundamental clinical skill, this is a structural weakness.
A diagnosis carries liability. If a tool suggests an erroneous conclusion and the doctor follows it, who answers? The practitioner, who did not exercise their judgement? The software vendor? The institution that deployed it? This is the void we described in our article on liability, with a vital stake here. As long as this question is unresolved, uses will remain confined to assistance rather than decision-making, and that is probably a good thing.
The rising issue: the patient who arrives with a diagnosis
A new phenomenon deserves to be flagged, because caregivers are encountering it more and more. Patients consult after querying an AI, sometimes with a specific hypothesis in mind.
This is not necessarily negative. An informed patient asks better questions, understands explanations better, and adheres more to treatment. The problem arises when the hypothesis provided by the machine becomes a conviction, or when it triggers disproportionate anxiety based on an answer produced without an examination.
Reasonable use, if one insists on consulting an AI about a health topic, comes down to one sentence: use it to understand the vocabulary and prepare your questions, never to form an opinion. This is not moralising, it is the direct consequence of what we know about the reliability of these tools.
What to take away
Medicine perfectly illustrates the gap between technical capability and real-world deployment. The capabilities are there, sometimes impressive. What holds things back is not the technology but everything else: clinical validation, robustness outside the lab, legal liability, integration into complex organisations.
This slowness annoys entrepreneurs and reassures patients, and both reactions are legitimate. In a field where an error cannot be fixed with an update, it is not absurd that adoption is slower than elsewhere. The real question is not about speeding up, it is about what we speed up: on administrative time given back to caregivers, very likely. On autonomous diagnosis, much less so.