The core argument rests on a straightforward observation: human oversight introduces latency, cognitive bias, and automation complacency. When a doctor reviews an AI recommendation, they don't evaluate it fresh β they anchor on the AI's output. Studies on automation bias show clinicians accept incorrect AI suggestions at alarming rates, especially under time pressure. Meanwhile, the AI doesn't get tired, doesn't have a bad day, and doesn't carry yesterday's diagnostic error into today's decision.
But here's the kicker the authors openly concede: virtually all the evidence comes from simulation studies. Retrospective chart reviews, standardized patient actors, vignette-based benchmarks. Not a single prospective trial in live patient care. The gap between "AI beats doctors on MIMIC-IV" and "AI safely manages a crashing septic patient in the ICU at 3 AM" is not a gap β it's a canyon.
The regulatory trap #
Current FDA guidance for clinical decision support software essentially requires a human in the loop for anything classified as high-risk. The European AI Act takes a similar approach. Both frameworks assume the human adds a safety layer. The JAMA piece argues this assumption is backwards once model performance exceeds a certain threshold β the human becomes the weak link.
This isn't theoretical. We've already seen it in radiology: AI-assisted radiologists catch more cancers than unaided radiologists, but fewer than AI alone on specific lesion types. The radiologist's "second look" sometimes overrides a correct AI call with a false negative. Multiply that across millions of scans and you get measurable harm.
What would autonomous deployment actually look like? #
The piece doesn't hand-wave this. It sketches a regulatory pathway: phased autonomy with defined escape hatches. Level 1 β AI suggests, human decides (current state). Level 2 β AI decides, human reviews retrospectively with audit trails. Level 3 β AI operates autonomously within a narrow, validated scope (e.g., diabetic retinopathy screening, already FDA-cleared for autonomous use via IDx-DR). Each level requires prospective evidence, not just benchmark scores.
The authors emphasize scope limitation. An autonomous AI for antibiotic stewardship in a hospital formulary is a different beast from an autonomous diagnostician for undifferentiated chest pain. Regulation should match the risk profile, not default to "human must approve."
The uncomfortable question #
If the evidence eventually shows autonomous AI kills fewer patients than human-AI teams for a defined task, does keeping a human in the loop become unethical? That's the question regulators will face within five years, not twenty. The JAMA piece is essentially telling them to start building the framework now β before the technology forces a reactive, restrictive response that locks in suboptimal care. The simulation-to-reality gap is real. But using it as a permanent excuse to mandate human oversight ignores the trajectory. We don't require a human to approve every autopilot adjustment on a 787. We required massive evidence, then we stepped back. Medicine deserves the same rigor β and the same willingness to follow evidence wherever it leads.
Next Parallel Exa Firecrawl lead Artificial Analysis Search Index for β
an AI side-hustle playbook, with plenty of directly applicable cases.