In April 2018, the FDA authorized IDx-DR, the first autonomous AI system permitted to diagnose a disease — more-than-mild diabetic retinopathy — without a physician reviewing the image. Its pivotal trial, published by Abràmoff and colleagues, reported 87.2% sensitivity and 90.7% specificity against a reference standard of retinal specialists (npj Digital Medicine, 2018). That is a genuinely strong result for a narrow, well-defined task. It is also the exception that proves how narrow the real advantage is.
The mechanism matters here. IDx-DR, like most cleared imaging AI, is a convolutional neural network trained on tens of thousands of labeled fundus photographs, optimized for one binary decision at one anatomical site, under one imaging protocol. It does not diagnose diabetes, does not read the rest of the eye exam, and was validated on a population deliberately similar to its training distribution. Radiology AI has followed the same template: as of 2026 the FDA has cleared more than 1,300 AI/ML-enabled devices, and roughly three-quarters of them sit in radiology — nodule detection, fracture triage, large-vessel-occlusion flagging, breast density scoring (IntuitionLabs FDA AI Medical Device Tracker). Nearly all are narrow single-task classifiers, not general diagnosticians.
Where these systems clearly beat unaided radiologists is triage speed and consistent recall on the specific pattern they were trained for — catching a large-vessel occlusion on CT in seconds to shorten stroke door-to-needle time, or flagging a missed lung nodule a fatigued reader skipped on read 40 of the day. Consistency, not superior perception, is the edge: the model never gets tired, never skips a case, and applies the identical decision boundary at 2am and 2pm.
Where it breaks down is deployment, not the lab. Google Health's own diabetic-retinopathy screening program is the clearest documented case. The pivotal validation looked excellent, but the prospective real-world rollout across nine primary-care sites in Thailand (Beede et al., CHI 2020; Ruamviboonsuk et al., Lancet Digital Health, 2022) found the algorithm rejected a meaningful share of images as too low-quality to grade — something rare in the curated trial dataset — forcing nurses to re-photograph patients or refer them anyway, disrupting the very workflow the tool was meant to speed up. Separately, a widely cited study by Zech et al. (PLOS Medicine, 2018) showed a pneumonia-detection CNN had partly learned hospital-specific scanner artifacts as a shortcut, and its accuracy dropped when tested on X-rays from a hospital it hadn't seen — a textbook case of dataset shift.
This is the pattern across imaging AI, not an isolated case: strong performance against a curated, single-site validation set, followed by a measurable drop — sometimes in accuracy, sometimes in usability, sometimes in clinician trust — once the tool meets the variability of a real population, real scanners, and real workflows. It is why clearance volume alone is a poor proxy for adoption — no public tracker yet reports what share of radiologists actually use these tools day to day, despite the volume of clearances (IntuitionLabs).
The count has kept climbing since: as of March 2026, the FDA's AI/ML-enabled device list has grown to roughly 1,450-1,500 cumulative clearances, with radiology still holding steady at about 76-77% of the total and the agency now clearing new algorithms at a pace of roughly 24-30 per month (IntuitionLabs FDA AI Medical Device Tracker; The Imaging Wire, March 2026). What has not changed is the deployment gap. A September 2025 UK rapid evidence assessment of autonomous chest x-ray classification found the same friction Google hit in Thailand years earlier — ambiguity in defining a truly "normal" image and unresolved regulatory liability under IR(ME)R — meaning no system has yet been given free rein to sign off studies without a radiologist in the loop. A parallel 2025 review of fracture-detection AI reached an identical verdict for commercial tools like BoneView and OsteoDetect: trial-grade accuracy, but real-world use remains supplementary, not autonomous.
Why this matters going forward: the FDA is beginning to require more external, multi-site validation and is piloting predetermined change control plans (PCCPs) so models can adapt post-market without a full new submission — an implicit admission that a single trial number was never going to be the whole story. The next competitive advantage in imaging AI will not come from squeezing another point of AUC out of a benchmark; it will come from whoever builds the most honest, most site-diverse evidence of what a model does after it leaves the lab.