The Microphone in the Exam Room

A decade of argument about machine learning in medicine was an argument about diagnosis, and the thing that reached the exam room is a microphone that drafts the visit note for a clinician to sign. This is about why, and the answer is four criteria in a statute passed in 2016 that decide what counts as a medical device. It is also an honest read of the evidence, which is eleven studies, mostly single-site and self-reported.

What Actually Showed Up

For most of the last decade the argument about machine learning in medicine was an argument about diagnosis. The model reads the scan. The model reads the retina. The model watches the ward overnight and tells you which patient is quietly getting worse. That was the promise, that is where the money and the conference programs went, and it is the version every clinician has been asked about at a dinner party. It is not what arrived. What arrived is a microphone.

An ambient documentation tool listens to a visit, transcribes it, and drafts the clinical note. The clinician reads the draft, edits it, and signs it. There is no diagnostic claim anywhere in that loop, no risk score, nothing that tells anybody what is wrong with the patient. It is close to the least glamorous thing a language model can do in a hospital, and it went from pilot to routine in about two years while the diagnostic models were still filing one submission at a time.

The usual explanations are that the scribes were simply better products, or that the documentation problem was more acute than the diagnostic one. I do not think either is the mechanism. The mechanism is a definition in a statute passed in December 2016, which decides whether a piece of software is a medical device, and the scribe was scoped to stay on the far side of it. Where this technology could land in medicine was settled by what counts as a device, and that has predicted adoption far better than accuracy ever did.

I want to be careful with that, because it reads as cynicism and it is not. The boundary is not a loophole somebody found. It is a considered line about which software a regulator can usefully inspect, and a note about a visit that already happened sits on the correct side of it. But a line drawn to manage risk also allocates capital, and the second effect turned out to be much larger than the first.

The Line That Decided It

Start with where the exclusion came from. Section 3060 of the 21st Century Cures Act, enacted 13 December 2016, amended section 520 of the Federal Food, Drug, and Cosmetic Act to exclude certain software functions from the definition of a device. The decision support carve-out is section 520(o)(1)(E), and it lists four criteria that must all be met. Fail one and the software is a device, with everything that follows: a marketing submission, a predicate or a De Novo, design controls, and a clearance number before anybody may sell it.

The first criterion disqualifies most of what people mean by clinical AI. The exclusion applies "unless the function is intended to acquire, process, or analyze a medical image or a signal from an in vitro diagnostic device or a pattern or signal from a signal acquisition system". Anything reading a scan, an ECG trace or a continuous glucose feed fails at the first hurdle and is a device by construction. FDA's September 2022 guidance even splits a single lab result, which is medical information, from continuous sampling of the same quantity, which is a signal. It shows in the numbers: counted by decision date, FDA's list of authorized AI-enabled devices is somewhere over twelve hundred entries and roughly three quarters radiology. That is not radiology being ahead. It is radiology having no other route.

The fourth criterion decides the interesting cases. The function must be "intended for the purpose of enabling" a clinician "to independently review the basis for the recommendations", so that it is not the intent that they rely primarily on it. FDA's guidance converts that into requirements: labeling must describe in plain language the approach used, the data relied upon so a clinician can assess whether it represents their own patients, and the validation results, explicitly including subgroups where performance is untested or highly variable.

The Clinician Signs It

A scribe is not intended to support or provide recommendations about prevention, diagnosis or treatment. It records what was said in a room and renders it as a chart entry. The same statute carries separate exclusions for software serving as electronic patient records and for administrative support of a health care facility, and drafting the note for a visit that already finished sits far closer to those. The four criteria are never reached, because they bind only once you are making the kind of claim the subsection is about.

Which is why "the clinician reviews and signs it" is doing an enormous amount of work in that sentence. It is the whole regulatory position and most of the liability position as well. The signature converts a machine-generated draft into a human attestation about a patient encounter, and everything downstream depends on it: the bill, the legal record, and what the next clinician reads three years from now when the patient turns up somewhere else.

An attestation is only a safeguard if the review is real, and I have found no published field measurement of how often a signed note contains something the encounter did not. The nearest thing is an instrument validation study from January 2025, concluding in the careful register these papers use that errors are present and must be evaluated to mitigate safety risk. That is a bench exercise rather than a measurement of practice, and the absence of the larger number is itself worth noticing.

What The Studies Measured

Documentation time moved, in minutes. A prospective quality improvement study followed 45 physicians across eight ambulatory disciplines at one academic center for three months. The tool was used in 9,629 of 17,428 encounters. Median time per note fell 0.57 minutes, which is nothing; median daily documentation time fell 6.89 minutes and total daily time in the record 19.95 minutes, which is a third of an evening. There is no control arm, the metrics come from the record system itself, and the authors flag wide variation between users.

The larger evaluation found the same shape and less certainty. Published on 1 May 2025, a hundred clinicians at a large California organization were measured three months either side of an April 2024 rollout. Time in notes per appointment fell from 6.2 to 5.3 minutes. Perceived mental demand and effort fell sharply. The share of clinicians meeting the burnout threshold fell from 42.1 to 35.1 percent and did not reach significance. That is the honest picture: workload perception moves hard, the clock moves by under a minute a visit, and burnout is a slower animal than either.

Here is what clinicians actually complain about. A qualitative study published in March 2025 interviewed twenty-two physicians from a pilot run over the winter of 2023 into 2024. They were positive on cognitive demand, work-life integration and, mostly, engagement with patients. They were largely negative on accuracy and style, specifically note length and how much editing was required. The tool is liked and it is not trusted to be left alone, and the second half is the important half.

And the field is smaller than it sounds. A systematic review published in April 2025 found eleven studies meeting its inclusion criteria, ten of them published in 2024, with a single product appearing in seven. Nine of ten reported some efficiency gain and seven of ten some wellness gain. Patient experience was assessed in three. Eleven studies is a literature at its beginning, and the correct posture toward it is interest rather than confidence.

Buying Back The Evening

Twenty minutes a day of a physician's evening is a real thing to buy. What is purchased is attention at the end of a clinic, and a willingness to still be doing the job in five years. The industry spent a decade insisting the prize was diagnostic performance and filing everything else under productivity, as though productivity were a consolation rather than the entire reason anyone automates anything.

The skeptical reading is that this is a productivity tool sold into a burnout crisis, and that the real disease is a documentation burden nobody with the power to cut it is going to cut. That is half right, and it does not survive an afternoon with people who use these tools. Both hold: the requirement is the illness, the scribe is symptomatic relief, and relief that returns evenings is worth paying for while the requirement stands. What would change my mind is a couple of years showing the returned time gets absorbed into more visits per session, at which point the clinician bought nothing and the schedule bought everything.

There is an inversion here worth naming. I have written before about why healthcare IT moves so slowly, and about how much of that slowness is earned rather than incompetent. Ambient documentation is the case that proves it from the other side. It moved fast precisely because it was scoped to avoid every part of the system that makes healthcare slow: no device submission, no diagnostic claim, no alteration to the clinical decision, and an output a human signs. It is a product designed around a regulatory boundary, and it went in at the speed of a subscription. That is not evidence the fast path is the good path. It says only that in a regulated market the cheapest thing to change is the claim you make.

Who It Works For

The transcription layer has a measured gap. Work published in April 2020 put five commercial speech recognition systems against structured interviews with 42 White and 73 Black speakers, 19.8 hours of audio matched on age and gender. Average word error rate was 0.35 for Black speakers against 0.19 for White speakers, and the gap held on identical phrases, which places it in the acoustic model rather than in vocabulary. I know of no equivalent published measurement for current clinical scribes, which is a hole in the evidence rather than a reason to assume it closed.

The reported failure has the same shape. The qualitative study above names limited functionality with patients who do not speak English as a barrier to adoption. Follow where that cost lands. A clinic whose panel is mostly English speaking gets the twenty minutes back; a clinic whose panel is not gets less of it, or none, while paying the same subscription. A tool whose benefit varies with the patient population is an equity question wearing the clothes of a procurement decision, and it will not show up in an aggregate time saving.

Bias enters through the label, not the model. Published in Science in October 2019: a widely used population health algorithm affecting millions of patients ranked people by predicted health care cost as a stand-in for illness. Because less is spent caring for Black patients at the same level of sickness, Black patients at a given score were considerably sicker. Correcting the proxy would have raised the share of Black patients flagged for extra help from 17.7 to 46.5 percent. Nothing about the model was broken. The thing it had been asked to predict was.

Eighteen Months From Now

A model degrades quietly, and the documented case is the Epic Sepsis Model, worth stating in full because it does more work than any argument I could make. An external validation published in June 2021 examined 27,697 patients across 38,455 hospitalizations at one academic health system. Area under the curve was 0.63. The model missed 1,709 of the 2,552 patients who developed sepsis, which is 67 percent of them, while generating alerts on 18 percent of all hospitalizations. It was already deployed at hundreds of United States hospitals before anyone outside the vendor validated it.

That is what I would put in front of anybody who has quietly concluded the device boundary is a quality boundary. It is not. It is a boundary about claims. A great deal of consequential software runs inside record systems on the non-device side of that line, deciding who gets looked at first, and the regulator's writ and the actual risk are not the same set. The boundary explains what gets reviewed. It says nothing about what works.

FDA's answer on the device side is the predetermined change control plan, finalized in December 2024, which lets a sponsor describe planned modifications and their validation up front so each retrain does not need a fresh submission; a draft guidance on lifecycle management followed on 7 January this year, its comment period closing in April. Both are aimed at devices. Everywhere else the monitoring plan is whatever the buyer wrote down, and the question to ask before signing is who checks this in eighteen months, on what data, against what baseline. I have rarely heard a good answer.

I should narrow the claim before I leave it. The boundary did not decide everything. Ambient documentation also arrived because transcription and summarization crossed a usefulness threshold at the same moment, and because the alternative was a human in the room or an hour of typing after dinner. Plenty of software that never went near the device definition failed anyway. What the boundary explains is not why this works, but why, of two roughly contemporaneous capabilities, this is the one that got to be routine first.

What stays with me is that the most consequential design review this technology will ever get happened in 2016, in a statute, before the thing existed. Four criteria about what counts as a device sorted an entire field into the part that ships and the part that files, and did it without naming a model, a method or a vendor. The microphone in the exam room is there because of a sentence written by people who were thinking about drug interaction alerts and spreadsheets, and it will still be there after the models behind it have been replaced twice.