AI Is Quietly Picking Your Doctor — And Lying About Why

AI Is Quietly Picking Your Doctor — And Lying About Why

Patients are asking ChatGPT which physician to see, and the AI is answering with confident, invisible bias. The model won’t tell you it’s doing it. The model doesn’t even know.

What happened

Researchers ran a prespecified randomized audit — 40,068 scored responses across seven large language models — testing what actually moves LLM physician recommendations when patients ask for help choosing a doctor. They built synthetic “physician cards” with independently randomized attributes (rating, fee, name, position) and ran them across 3,024 choice sets, nine prompt paraphrases, and three patient personas. The dominant finding is blunt: reputation and price swamp everything else. A rating jump from 3.9 to 4.7 raised selection probability by 31.4 percentage points; a fee increase from $90 to $190 cut it by 20.0 pp. Demographic effects exist but run opposite to what human hiring-audit literature predicts — female-signaled names gained 2.5 pp, and Hispanic-, South-Asian-, and Black-signaled names gained 1.3–2.9 pp over White-signaled names, worth roughly $7–$14 per visit in fee-equivalent terms. A prompt-level ranking artifact also showed up: being listed first in the prompt — with no other difference — was worth $11 in fee-equivalent advantage. Crucially, models cited gender or ethnicity as a reason in at most 0.03% of stated explanations, meaning AI-generated content disclosure and self-report transparency mechanisms would detect essentially none of this.

Cold read

This is a well-designed audit, but “synthetic physician cards” is doing enormous load-bearing work — real clinical decisions layer in insurance networks, appointment availability, specialty depth, and patient history that this design deliberately strips out. The demographic tilts found here are small in absolute terms and run counter to discrimination narratives, which makes them interesting but also easy to weaponize selectively; a 2.5 pp lift for female names does not mean the system is “fair,” it means the bias is just differently shaped than expected. The study covers seven models at a single point in time with frozen stimuli — benchmark contamination risk is low by design, but model behavior drifts with every update, so today’s audit is yesterday’s news by next quarter. One reasoning model failed the “auditability gate” entirely, which is buried in the abstract but is arguably the most alarming finding: a class of models can’t even be audited by this method. Finally, the audit measures what models recommend, not what patients actually do — conversion from LLM suggestion to booked appointment is an unmeasured and potentially large gap.

What it means for you

  • Signal maturity: 3/5 — rigorous method, but synthetic stimuli limits real-world generalizability
  • Who gets hurt: Healthcare marketplace operators (Zocdoc-style platforms, insurer directories) whose physician AI visibility and share of voice they cannot currently audit or guarantee
  • What breaks if this is true: Any compliance or ethics framework that relies on model self-explanation to detect demographic bias is structurally blind — regulatory “explain your AI” mandates become theater
  • Why it might not land: The demographic effects are small and directionally counterintuitive, giving platforms and regulators easy political cover to deprioritize; the position-bias and rating-dominance findings are more actionable but less headline-grabby
  • Watch for: EU AI Act enforcement bodies or FTC issuing guidance that explicitly rejects self-reported model explanations as sufficient audit evidence — that would be the regulatory tripwire that makes this paper matter operationally

Forecast as of 2026-08-17

By Q3 2027, at least one major healthcare platform or insurer will publicly commission or cite a behavioral algorithm audit (not a self-report disclosure) of their LLM recommendation layer, driven by regulatory pressure rather than voluntary transparency — the methodology in this paper is replicable enough to become the template.


Source: Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice — Syeda Anshrah Gillani, Mirza Samad Ahmed Baig. https://arxiv.org/abs/2608.14399v1

Similar Posts