AI Scribes Are Quietly Fabricating Medical Records at Scale

AI Scribes Are Quietly Fabricating Medical Records at Scale

One in three clinical notes generated by commercial AI scribes contains a verified error. Not a typo. Not a stylistic quibble. A failure — in allergy documentation, medication data, or invented patient identity — that a signing clinician is supposed to catch but statistically often won’t.

What happened

Researchers audited three commercial AI scribes against 142 real UK and US clinical consultations, generating 565 notes. A multi-stage pipeline produced 13,678 candidate errors; after importance filtering, 5,898 went to an adversarial LLM-as-judge panel — two models from different families, each tasked with refuting findings — and 618 survived. The headline number: 31.3% of notes [CI: 27.0–35.6%] carried a verified failure. Stripping out two error classes that patient records would have prefilled (invented identity and dates), the rate holds at 24.8% [20.8–29.0%]. The failure modes aren’t trivial: allergy and medication errors, hallucination of patient identity, and physical examination findings written into notes for telephone consultations where no exam occurred. One failure mode stood entirely outside standard taxonomies: a treatment the clinician verbally retracted, recorded in the note as delivered care. Clinician adjudication was rigorous — a physician author upheld 20 of 21 sampled findings (95.2%), and an independent clinician upheld 12 of 12.

Cold read

The adversarial review pipeline is the paper’s most important finding — and its biggest liability. The authors themselves demonstrate that prompt engineering alone moves the verified failure rate from 9.3% to 79.0%, and the choice of reviewing model family shifts it from 27.8% to 54.8% at the same instruction. That’s not a minor calibration issue; it means the “31.3%” number is instrument-dependent in a way the headline doesn’t acknowledge. The authors are admirably honest about this — they flag that published audit studies disagree by margins that instrument differences alone can explain — but founders and procurement teams won’t read the methodology section. This is also a sample of 142 consultations across three products; we don’t know which products, in what ratio, or how representative the “authored scenarios” are of real-world deployment volume. The finding that omission comprises only 23.1% of errors (vs. 54–86% in prior audits) likely reflects the adversarial filter design as much as any truth about these specific scribes. Clinician sign-off is presented throughout the industry as the safety net; this paper doesn’t measure how often clinicians actually catch errors before signing, which is the number that would determine real-world patient harm.

What it means for you

  • Signal maturity: 3/5 — Rigorous methodology, but instrument sensitivity limits external validity
  • Who gets hurt: Health-tech founders selling ambient scribe features as “clinician-reviewed” to enterprise health systems, and the compliance officers who signed off on that framing
  • What breaks if this is true: The liability shield of “a clinician signs every note” collapses if clinicians are effectively rubber-stamping a 1-in-4 error rate; malpractice exposure lands on the practice, not the vendor
  • Why it might not land: Three unnamed products, 142 consultations, and an openly instrument-sensitive pipeline give defense lawyers and enterprise procurement teams plenty of room to wait for replication at scale
  • Watch for: A US health system or NHS trust citing this paper in a vendor contract renegotiation or a CMS/MHRA regulatory inquiry into ambient scribe certification standards

Forecast as of 2026-09-01

By Q2 2027, at least one major ambient scribe vendor (Nuance/DAX, Abridge, or Nabla) will publish a formal rebuttal study or commission an independent audit explicitly referencing this paper’s methodology — and the error rate they report will be below 10%, illustrating exactly the instrument-sensitivity problem the authors described.


Source: One note in three: a verified census of three deployed AI scribes, and the instrument that counted it — Sebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris. https://arxiv.org/abs/2608.31017v1

Similar Posts