Your AI Radiologist Aces the Test, Then Misses the Tumor
Your AI Radiologist Aces the Test, Then Misses the Tumor
A new paper proves that the more precisely you task an AI model, the blinder it becomes to everything else. That’s not a niche research curiosity — that’s a product liability argument waiting to happen in every vertical AI deployment on the planet.
What happened

Researcher Kwan Soo Shin ran a systematic study across radiology text scenarios, driving text scenarios, and chest-radiograph vision tasks to test a simple, damning hypothesis: do narrowly tasked models suppress detection of safety-critical signals they can report when unconstrained? The answer is yes, universally. Every model tested showed this “Inattentional Gap” — the same model that catches a hazard when asked broadly will miss it when focused on a different task. Critically, suppression did not diminish with scale, meaning your expensive frontier model is not buying you out of this problem. The effect also persisted in a reasoning model — the chain-of-thought architecture many founders are betting on as the path to reliable AI agents — and varied more by model family than by size. The paper argues this decouples benchmark contamination-style measured safety from real-world safety: a system can score near-perfectly on specified hazards while remaining blind to the ones that actually cause harm.
Cold read
This is a single-author paper, and the abstract offers no sample sizes, no confidence intervals, and no named models — making independent verification impossible from the outside. “Every model tested” is a strong claim whose weight depends entirely on which models were tested and how many, information not surfaced in the abstract. The radiology and driving scenarios are plausibly high-stakes, but how well these text-based and chest-radiograph tasks generalize to, say, a contract-review AI or a customer-support agent is an open question the paper does not appear to answer. The mechanism the paper proposes — that this is analogous to human inattentional blindness but arises from a different mechanism — is interesting but also conveniently unfalsifiable at the abstract level; without knowing the proposed mechanism, you cannot design around it. The finding that variance tracks model family more than size is useful directionally, but without naming families, it’s actionable for nobody.
What it means for you
- Signal maturity: 3/5 — Real phenomenon, thin replication package so far
- Who gets hurt: Vertical AI builders shipping narrow-task copilots into high-stakes environments: medical, legal, industrial safety, autonomous vehicle stacks
- What breaks if this is true: Your model card and benchmark numbers are measuring the wrong thing — compliance and procurement teams will eventually figure this out, and your liability exposure is larger than your eval suite suggests
- Why it might not land: Most enterprise AI deployments are already scoped narrowly by design, and “we only guarantee what we were tested on” is a defensible contractual position; regulators may accept benchmark evidence longer than they should
- Watch for: The first successful product-liability or negligence case citing AI task-conditioning as a causal factor in a missed safety signal — that’s the moment this paper gets cited in every courtroom and procurement RFP
Forecast as of 2026-06-26
By Q2 2027, at least one major AI safety benchmark consortium (MLCommons, METR, or equivalent) will add an explicit “unconstrained vs. task-conditioned” detection delta metric to their evaluation suites, directly citing the Inattentional Gap framing — or the concept will remain confined to academic discourse with zero procurement impact.
Source: The Inattentional Gap: Task-Conditioned Language and Vision Models Omit the Safety-Critical Signals They Can Otherwise Report — Kwan Soo Shin. https://arxiv.org/abs/2606.26529v1
