AI Is Referring Your Patients to Doctors Who Don’t Exist

AI Is Referring Your Patients to Doctors Who Don’t Exist

You asked an AI which cardiologist to trust. It gave you a name, a city, a confident answer — and a 96% chance the doctor isn’t real. This isn’t a fringe bug. It’s the default behavior of every AI assistant not wired to search.

What happened

Ibrahim and Zaki audited AI provider recommendations across four regulated service domains — Medicare clinicians, medical facilities, and SEC-registered investment advisers — in all 100 largest U.S. metro areas, then matched every recommendation against official government registries. They tested three configurations: an open-weight model, a proprietary model without web search, and the same proprietary model with search enabled. Without search, the results are damning: only 4% of the open-weight model’s recommended doctors match a real clinician in the queried city, and even those matches are coincidences — the “matched” doctors are no more likely to be primary-care physicians than a random name pulled from the registry. That is hallucination operating at near-total scale. The proprietary model without search fares marginally better at 11%, but the more alarming finding is what the models recommend when names are wrong: advisory firms recommended without search carry SEC misconduct disclosures at 3.6× the registry base rate. Enabling retrieval-augmented generation via search lifts real-provider match rates to 64–71% and flips the misconduct signal to below the base rate. In restaurants — where visibility and quality can be measured separately — search-enabled models show a 3–5× review-count premium but a rating premium of at most 0.1 stars, suggesting AI visibility is correlated with noise volume, not quality. Without search, real recommendations also concentrate disproportionately in the largest metros; search largely eliminates this geographic penalty.

Cold read

This is an audit, not a causal experiment — the paper cannot tell you why misconduct-heavy firms appear more often without search, only that they do; the mechanism (training data skew, name salience, PR spend) is unresolved. The “4% match” figure covers only whether the name exists in the registry in that city, not whether the recommendation is clinically appropriate or harmful in practice — it measures fabrication rate, not harm rate. The restaurant finding is correlational and doesn’t prove that high-share-of-voice businesses are gaming anything; review count is a confounder for genuine popularity. The 64–71% match rate with search sounds like progress but still means roughly one-in-three recommendations in thin-coverage domains is unverifiable — the paper doesn’t break out which domains drive the remaining failures. Finally, the study tests specific model configurations as of a snapshot in time; model updates and search integrations are moving fast enough that these numbers could age out within quarters.

What it means for you

  • Signal maturity: 4/5 — Large-scale registry-matched audit with specific numbers; methodology is reproducible and hard to dismiss
  • Who gets hurt: Healthcare directories, financial adviser platforms, and any local-services marketplace that competes with AI-generated referrals — your verified listings are being bypassed by confident fabrications
  • What breaks if this is true: The implicit trust model behind AI assistants — that a confident, cited answer reflects a real, vetted provider — is false by default without search, and users have no way to detect it because the output “carries no sign that its recommendations were never verified
  • Why it might not land: Most AI consumer products are already moving toward search-grounded defaults; vendors will argue this paper describes yesterday’s problem and point to improving match rates as evidence
  • Watch for: Regulatory action from CMS or SEC requiring AI assistants to disclose retrieval configuration and verification status when making provider referrals — the misconduct-multiplier finding is exactly the kind of data that moves regulators

Forecast as of 2026-09-17

By Q3 2027, at least one major AI assistant provider (OpenAI, Google, or Anthropic) will publish a policy or technical disclosure specifically addressing provider-recommendation verification — either a retrieval-configuration disclosure requirement or a domain-specific refusal to recommend unverified clinicians and advisers — driven partly by this class of audit research reaching regulatory audiences.


Source: Understanding AI Provider Recommendations in Local Service Markets — Hazem Ibrahim, Yasir Zaki. https://arxiv.org/abs/2609.18341v1

Similar Posts