Evidence

AI in medicine: what the literature says

The real question is not "should there be AI in medicine". It is: well bounded, does it help the physician? The recent literature answers yes, on three fronts.

Our position

Well bounded, AI helps the physician

Administrative support, cognitive load, differential diagnosis. Here is the data, and what the tool does with it.

Our line fits in one sentence: the AI suggests, the physician decides. The pages that follow show what research establishes, and where it cautions, hiding nothing.

Front 1

Time given back to the patient

Nearly 16,000 documentation hours given back to physicians, in a single network.2

The largest published deployment: more than 7,200 physicians, 2.5 million encounters, the equivalent of nearly 1,800 workdays saved.1217 The first randomized trial confirms the drop in time spent in the note.3 And within the first weeks, sites report marked burnout drops, up to 63%.4

What AtomicMD does with it: the heart of the product. One dictation, and the note, the orders, the Quebec forms and the billing prefill. The administrative work, lightened in one command. You verify.

Front 2

Lighten the load, catch the red flags

In the emergency department, about one patient in eighteen leaves with a diagnostic error.20

The cause is almost never ignorance. It is the load: interruptions, fatigue, working-memory overload. The dangerous missed diagnoses are often known presentations, seen at the wrong moment.

What AtomicMD does with it: it offloads working memory. It offers the questions that change management, recalls reassessments, sorts the dashboard by the next useful step. Less load, fewer misses.

Front 3

A second look at the differential

In Nature, in a randomized double-blind study: a diagnostic AI produced better differentials than front-line physicians.7

159 clinical scenarios, validated simulated patients, 20 physicians as the comparison: the AI's differential was more accurate at every level measured.7 And in the emergency department, 2026 multicentre work shows the best models catch the large majority of missed diagnostic opportunities in flagged charts.8

What AtomicMD does with it: that second look, in service of yours. Not a chat window: a differential integrated into the note, at the right moment, worst-first (common, serious, rare but catastrophic), so the dangerous is never missed. You keep the hand, and the last word.

The real objection

"But AI hallucinates"

It is the most common criticism, and it is not wrong. A model left alone can invent a credible detail, including in transcription.910 Misused, AI can also introduce a bias or dull a reflex.1112

The literature does not conclude that AI must be abandoned. It concludes that the physician must stay in control. That is also what the WHO, the Collège des médecins and the American Medical Association say, the latter speaking of "augmented intelligence": the AI assists, the human decides.131618

This is the constitution of AtomicMD: nothing inserts without your click, every fact carries its source, every score variable carries the exact quote from your dictation and nothing is computed for you, and a checker compares the note to your dictation.

The debate is not AI or not. It is well-bounded AI. That is what AtomicMD is.

Sources

References

Verified selection (2023 to 2026). Links revalidated before publication.

Administrative support

1Tierney AA et al. "Ambient Artificial Intelligence Scribes to Alleviate the Burden of Clinical Documentation." NEJM Catalyst, 2024. doi:10.1056/CAT.23.0404
2Tierney AA et al. "Ambient AI Scribes: Learnings after 1 Year and over 2.5 Million Uses." NEJM Catalyst, 2025. doi:10.1056/CAT.25.0040
3Lukac PJ et al. "Ambient AI Scribes in Clinical Practice: A Randomized Trial." NEJM AI, 2025. doi:10.1056/AIoa2501000
4Peterson Health Technology Institute. "Ambient Scribes: Evaluation." Report, 2025.
17Tierney AA et al. Large-scale deployment of an ambient AI scribe (Kaiser Permanente). NEJM Catalyst, 2024.

Cognitive load and diagnostic safety

20AHRQ. "Diagnostic Errors in the Emergency Department: A Systematic Review." 2022.

Differential diagnosis

5Kanjee Z et al. "Accuracy of a Generative AI Model in a Complex Diagnostic Challenge." JAMA, 2023.
6Goh E et al. "Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial." JAMA Network Open, 2024.
7Tu T et al. "Towards conversational diagnostic artificial intelligence." Nature, 2025;642:442-450. doi:10.1038/s41586-025-08866-7
8LLM detection of missed diagnostic opportunities in the emergency department. Multicenter diagnostic study (288 encounters, 9 hospitals), 2026.

Limits, oversight and official positions

9Omar M et al. "Multi-model assurance analysis of hallucinations in medical LLMs." Communications Medicine, 2025. doi:10.1038/s43856-025-01021-3
10Automatic speech transcription hallucinations. Proceedings of ACM FAccT, 2024. doi:10.1145/3630106.3658996
11Qazi H et al. "Effect of flawed AI recommendations on physician diagnostic reasoning." NEJM AI, 2026. doi:10.1056/AIoa2501001
12De-skilling of endoscopists after exposure to AI. The Lancet Gastroenterology & Hepatology, 2025;10(10):896-903.
13World Health Organization. "Ethics and governance of AI for health: Large multi-modal models." 2024.
14Health Canada. "Software as a Medical Device (SaMD): definition and classification." 2019.
15FDA. "Clinical Decision Support Software." Final guidance, 2026.
16Collège des médecins du Québec. "AI and the practice of medicine."
18American Medical Association. "Augmented intelligence in medicine."
19Law 25 (Québec), decision based exclusively on automated processing.