A decade of argument about machine learning in medicine was an argument about diagnosis, and the thing that reached the exam room is a microphone that drafts the visit note for a clinician to sign. That swap is what this piece is about, and the answer is four criteria in a statute passed in 2016 that decide what counts as a medical device. The evidence underneath is eleven studies, mostly single-site and self-reported.
12 June 2025·8 min read·aihealthcare
For most of the last decade the argument about machine learning in medicine was an argument about diagnosis. The model reads the scan. The model reads the retina. The model watches the ward overnight and tells you which patient is quietly getting worse. That was the promise, that is where the money and the conference programs went, and it is the version every clinician has been asked about at a dinner party. What arrived is a microphone.
An ambient documentation tool listens to a visit, transcribes it, and drafts the clinical note. The clinician reads the draft, edits it, and signs it. Not one step in that loop makes a diagnostic claim, scores a risk, or tells anybody what is wrong with the patient. Drafting a note is close to the least glamorous thing a language model can do in a hospital, and it went from pilot to routine in about two years while the diagnostic models were still filing one submission at a time.
The usual explanations are that the scribes were simply better products, or that the documentation problem was more acute than the diagnostic one. I do not think either is the mechanism. The mechanism is a definition in a statute passed in December 2016, which decides whether a piece of software is a medical device, and the scribe was scoped to stay on the far side of it. Where this technology could land in medicine was settled by what counts as a device, and that has predicted adoption far better than accuracy ever did.
This reads as cynicism and it is not. The boundary is a considered line about which software a regulator can usefully inspect, and a note about a visit that already happened sits on the correct side of it. But a line drawn to manage risk also allocates capital, and the second effect turned out to be much larger than the first.
Start with where the exclusion came from. Section 3060 of the 21st Century Cures Act, enacted 13 December 2016, amended section 520 of the Federal Food, Drug, and Cosmetic Act to exclude certain software functions from the definition of a device. The decision support carve-out is section 520(o)(1)(E), and it lists four criteria that must all be met. Fail one and the software is a device. Everything follows from that: a marketing submission, a predicate or a De Novo, design controls, and a clearance number before anybody may sell it.
The first criterion disqualifies most of what people mean by clinical AI. The exclusion applies "unless the function is intended to acquire, process, or analyze a medical image or a signal from an in vitro diagnostic device or a pattern or signal from a signal acquisition system". Anything reading a scan, an ECG trace, or a continuous glucose feed fails at the first hurdle and is a device by construction. FDA's September 2022 guidance even splits a single lab result, which is medical information, from continuous sampling of the same quantity, which is a signal. It shows in the numbers: counted by decision date, FDA's list of authorized AI-enabled devices is somewhere over twelve hundred entries and roughly three quarters radiology. Radiology dominates that list because it has no other route.
The fourth criterion decides the cases the first three do not. The function must be "intended for the purpose of enabling" a clinician "to independently review the basis for the recommendations", so that it is not the intent that they rely primarily on it. FDA's guidance converts that into requirements: labeling must describe in plain language the approach used, the data relied upon so a clinician can assess whether it represents their own patients, and the validation results, explicitly including subgroups where performance is untested or highly variable.
A scribe is not intended to support or provide recommendations about prevention, diagnosis, or treatment. It records what was said in a room and renders it as a chart entry. The same statute carries separate exclusions for software serving as electronic patient records and for administrative support of a health care facility, and drafting the note for a visit that already finished sits far closer to those. The four criteria are never reached, because they bind only once you are making the kind of claim the subsection is about.
"The clinician reviews and signs it" is doing an enormous amount of work. The review clause is the whole regulatory position and most of the liability position as well. The signature converts a machine-generated draft into a human attestation about a patient encounter, and everything downstream depends on it: the bill, the legal record, and what the next clinician reads three years from now when the patient turns up somewhere else.
An attestation is only a safeguard if the review is real, and I have found no published field measurement of how often a signed note contains something the encounter did not. The nearest thing is an instrument validation study from January 2025, concluding in the careful register these papers use that errors are present and must be evaluated to mitigate safety risk. That study is a bench exercise. Nobody has measured practice, and the absence of the larger number is itself worth noticing.
Documentation time moved, in minutes. A prospective quality improvement study followed 45 physicians across eight ambulatory disciplines at one academic center for three months. The tool was used in 9,629 of 17,428 encounters. Median time per note fell 0.57 minutes, which is nothing; median daily documentation time fell 6.89 minutes and total daily time in the record 19.95 minutes, which is a third of an evening. The study has no control arm, the metrics come from the record system itself, and the authors flag wide variation between users.
The larger evaluation found the same shape and less certainty. Published on 1 May 2025, a hundred clinicians at a large California organization were measured three months either side of an April 2024 rollout. Time in notes per appointment fell from 6.2 to 5.3 minutes. Perceived mental demand and effort fell sharply. The share of clinicians meeting the burnout threshold fell from 42.1 to 35.1 percent and did not reach significance. Those numbers are the picture: workload perception moves hard, the clock moves by under a minute a visit, and burnout is a slower animal than either.
Here is what clinicians actually complain about. A qualitative study published in March 2025 interviewed twenty-two physicians from a pilot run over the winter of 2023 into 2024. They were positive on cognitive demand, work-life integration, and, mostly, engagement with patients. They were largely negative on accuracy and style, specifically note length and how much editing was required. The tool is liked. It is not trusted to be left alone, and the second half is the important half.
And the field is smaller than it sounds. A systematic review published in April 2025 found eleven studies meeting its inclusion criteria, ten of them published in 2024, with a single product appearing in seven. Nine of ten reported some efficiency gain and seven of ten some wellness gain. Patient experience was assessed in three. Eleven studies is a literature at its beginning, and it earns interest without yet earning confidence.
Twenty minutes a day of a physician's evening is a real thing to buy. What is purchased is attention at the end of a clinic, and a willingness to still be doing the job in five years. The industry spent a decade insisting the prize was diagnostic performance and filing everything else under productivity, as though productivity were a consolation prize, when it is the entire reason anyone automates anything.
The skeptical reading is that this is a productivity tool sold into a burnout crisis, and that the real disease is a documentation burden nobody with the power to cut it is going to cut. The skepticism is half right. It does not survive an afternoon with people who use these tools. Both hold: the requirement is the illness, the scribe is symptomatic relief, and relief that returns evenings is worth paying for while the requirement stands. What would change my mind is a couple of years showing the returned time gets absorbed into more visits per session, at which point the clinician bought nothing and the schedule bought everything.
Something inverts here, and it is worth naming. I have written before about why healthcare IT moves so slowly, and about how much of that slowness is earned rather than incompetent. Ambient documentation is the case that proves it from the other side. It moved fast precisely because it was scoped to avoid every part of the system that makes healthcare slow: no device submission, no diagnostic claim, no alteration to the clinical decision, and an output a human signs. The scribe is a product designed around a regulatory boundary, and it went in at the speed of a subscription. The speed is not evidence the fast path is the good path.
The transcription layer has a measured gap. Work published in April 2020 put five commercial speech recognition systems against structured interviews with 42 White and 73 Black speakers, 19.8 hours of audio matched on age and gender. Average word error rate was 0.35 for Black speakers against 0.19 for White speakers. The gap held on identical phrases. That points at the acoustic model, since the words were identical. I know of no equivalent published measurement for current clinical scribes, which leaves a hole in the evidence and no reason to assume it closed.
The reported failure has the same shape. The qualitative study above names limited functionality with patients who do not speak English as a barrier to adoption. Follow where that cost lands. A clinic whose panel is mostly English speaking gets the twenty minutes back; a clinic whose panel is not gets less of it, or none, while paying the same subscription. A tool whose benefit varies with the patient population is an equity question wearing the clothes of a procurement decision, and it will not show up in an aggregate time saving.
The label is where bias enters, and the model only inherits it. Published in Science in October 2019: a widely used population health algorithm affecting millions of patients ranked people by predicted health care cost as a stand-in for illness. Because less is spent caring for Black patients at the same level of sickness, Black patients at a given score were considerably sicker. Correcting the proxy would have raised the share of Black patients flagged for extra help from 17.7 to 46.5 percent. The model itself was fine. The broken part was what it had been asked to predict.
A model degrades quietly, and the documented case is the Epic Sepsis Model, worth stating in full because it does more work than any argument I could make. An external validation published in June 2021 examined 27,697 patients across 38,455 hospitalizations at one academic health system. Area under the curve was 0.63. The model missed 1,709 of the 2,552 patients who developed sepsis, which is 67 percent of them, while generating alerts on 18 percent of all hospitalizations. It was already deployed at hundreds of United States hospitals before anyone outside the vendor validated it.
That is what I would put in front of anybody who has quietly concluded the device boundary is a quality boundary. The device boundary marks what a product may claim, and quality is a separate question. A great deal of consequential software runs inside record systems on the non-device side of that line, deciding who gets looked at first, and the regulator's writ and the risk are not the same set. The boundary explains what gets reviewed. It says nothing about what works.
FDA's answer on the device side is the predetermined change control plan, finalized in December 2024, which lets a sponsor describe planned modifications and their validation up front so each retrain does not need a fresh submission; a draft guidance on lifecycle management followed on 7 January this year, its comment period closing in April. Both are aimed at devices. Everywhere else the monitoring plan is whatever the buyer wrote down, and the question to ask before signing is who checks this in eighteen months, on what data, against what baseline. I have rarely heard a good answer.
The boundary did not decide everything. Ambient documentation also arrived because transcription and summarization crossed a usefulness threshold at the same moment, and because the alternative was a human in the room or an hour of typing after dinner. Plenty of software that never went near the device definition failed anyway. What the boundary explains is why, of two roughly contemporaneous capabilities, this is the one that got to be routine first.
The most consequential design review this technology will ever get happened in 2016. In a statute, before the thing existed. Four criteria about what counts as a device sorted an entire field into the part that ships and the part that files, and did it without naming a model, a method, or a vendor. That microphone is there because of a sentence written by people who were thinking about drug interaction alerts and spreadsheets, and it will still be there after the models behind it have been replaced twice.