FDA has now written down how it might regulate generative AI medical devices, and the shape of it is the one I have been arguing for since my first foundation-model pre-submission in 2024. On August 18, 2026, CDRH's Digital Health Center of Excellence released Considerations for the Regulation of Generative AI-Enabled Medical Devices, a discussion paper and request for feedback under docket FDA-2026-N-7874. Comments are due October 19, 2026. It is not guidance, and it says so on page one, but it is still the most concrete thing FDA has published about LLM-enabled devices.
Three ideas carry the paper. Risk gets scored on two axes: how independently the function acts, and how bad a wrong output would be. Premarket evaluation becomes competency-based, "inspired, at a high level, by how human clinicians are evaluated and credentialed": benchmarking first, then clinical confirmation in progressively realistic settings. And FDA asks whether it should accept more premarket uncertainty in exchange for continuous postmarket monitoring. That last question is the one I put to the Digital Health Advisory Committee in person in November 2024, on FDA's own transcript, when I asked the agency to turn postmarket surveillance from a stick into a carrot.
That DHAC exchange was not a one-off. In July 2024 I published a foundation model FAQ that flagged benchmark contamination before FDA gave it a section number. In October I published a pre-submission that asked FDA, in writing, whether an independent LLM could adjudicate a test set and whether nightly re-benchmarking would catch a cloud vendor silently changing the model. In November I argued the postmarket trade in person at DHAC, and in December, from the RSNA podium, I told sponsors to constrain the shell so a user cannot turn a medical device into ChatGPT. FDA might be reading my thought leadership pieces all along :).
The paper is 30 pages with 26 numbered discussion questions, written by DHCoE under Director Rick Abramson and CDRH Director Michelle Tarver, with Acting Commissioner Kyle Diamantas framing it as part of the push to move AI medical products to market faster. It grows out of the November 2024 DHAC meeting on generative AI, which the paper cites as its origin.
Read the disclaimer before the hype. The paper "does not represent draft or final guidance," is "not intended to propose or implement policy changes," and does not address whether the approaches fit within FDA's existing legal authority. RAPS's same-day roundup called it draft guidance. It is not. It is an early look at the evidence structure FDA is considering, offered before the agency commits, so the comment period is where that structure gets negotiated. FDA also restates the principle underneath everything: it "does not regulate GenAI as such; it regulates medical devices."
FDA's two-axis framework scores a generative AI function by what it does and by the consequence of relying on an incorrect output. The activity axis runs from non-directive information (a cardiovascular risk score) through action-directing information (a strongly worded push to seek emergency care) to action-taking under clinician supervision, and finally to fully autonomous action. The consequences axis runs from limited to severe. Risk rises from lower left to upper right.
FDA's examples show where its attention is. Hydrocortisone for poison ivy sits low. Whether to go to the ED for chest pain, or how to adjust basal insulin, sits high even though both are "just information." Autonomously prescribing antibiotics for confirmed strep is lower than autonomously starting a thrombolytic order set for stroke. The evidence expectation follows the position on the grid rather than the presence of an LLM.
Section IV's finer points will reshape intended-use statements. Directiveness is a continuum judged on substance, and adding "talk to your doctor" does not make an action-directing output less directive. Patient-facing functions may move up the consequences axis because patients cannot check the output. Multi-turn conversations are assessed on realistic trajectories, since a chat that starts informational can drift into directing action. And escalation is scored in both directions: sending everyone to the ED counts as harm.
Your intended-use sentence now has to say how independently the function acts and how bad a wrong answer is, and your risk file has to defend both placements. That is the line I drew in my UpDoc analysis in June. FDA had cleared conversation on the outside and a deterministic protocol on the inside, not an autonomous LLM physician.
The grid does not replace the January 2026 revisions to the clinical decision support and general wellness guidances. Those two documents draw the outer boundary of what is a device; the paper grades what sits inside it. Section III restates the sorting: many software functions are not devices under section 520(o), others sit under enforcement discretion, and FDA focuses on device functions whose failure could hurt a patient.
The revised CDS guidance says patient-facing recommendations are devices; the paper pushes patient-facing functions up the consequences axis for the same reason, and footnote 13 ties outputs the user cannot evaluate to "the independent-review criterion of the clinical decision support exclusion under section 520(o)(1)(E)." The CDS guidance now tolerates a single recommendation when only one option is clinically appropriate; the paper's non-directive-to-action-directing continuum is that idea with a risk gradient attached. The wellness lane, widened in January to non-invasive wearables estimating blood pressure and glucose, does not stretch to a chatbot that directs care, because a "talk to your doctor" line does not reduce directiveness. Multi-turn drift is the failure mode to watch: a product can start in the wellness or non-device CDS lane and talk its way into the device column.
FDA's competency-based approach evaluates a generative AI device the way medicine evaluates a physician: standardized benchmarking of knowledge, safety behavior, and generalizability, then clinical confirmation in progressively realistic settings, then ongoing assessment in use. Clinicians are not tested on every scenario they might meet, FDA reasons, and neither can a device with open-ended inputs be. The paper cites Patel and Blumenthal (JAMA Health Forum, 2026), Bergman, Wachter, and Emanuel's licensure framework (JAMA, 2026), and Freyer et al. (Nature Medicine, 2025), then adapts the idea to device law.
Two sentences in Section V decide your architecture. The evaluation target is "the final user-facing device, as configured and intended to be deployed for real-world use, and not the foundation model standing alone." And the approach is "proportionate to its risk," with the two-axis grid setting how much benchmarking and confirmation.
Benchmarking is FDA's word for high-throughput, non-clinical testing of the deployed configuration. The paper proposes ten elements, chosen per device by intended use and risk:
| Group | Element | What FDA wants probed (Appendix A) |
|---|---|---|
| Safety | S.1 Safety-critical recognition and escalation | Time to escalation in evolving encounters; resistance to over-reassurance; under- and over-escalation |
| Safety | S.2 Scope maintenance and boundary adherence | Adversarial prompting, prompt injection, emotional manipulation, multi-turn drift out of scope; over-refusal also fails |
| Safety | S.3 Calibration, uncertainty, clinical deferral | False confidence on contested or outdated information is a safety failure |
| Clinical proficiency | E.1 Clinical knowledge and task fidelity | Current guidelines, contraindications, interactions, special populations |
| Clinical proficiency | E.2 Information gathering and analysis | Differential completeness, follow-up questions, premature closure |
| Clinical proficiency | E.3 Quantitative and measurement analysis | Weight- and renal-based dosing, unit conversions, implausible values |
| Clinical proficiency | E.4 Communication and comprehension | Health literacy, empathy, coercive language, automation bias |
| Generalizability | R.1 Robustness, reliability, reproducibility | Repeated runs, paraphrases, input order, long conversations |
| Generalizability | R.2 Subgroup performance | Demographics, dialects, accents, literacy levels |
| Agentic | A.1 Agentic competencies | Planning inside the safety envelope, tool-error recognition, human checkpoints before irreversible actions |
The method principles will be familiar to anyone who has run a reader study: prespecified methods and acceptance criteria, rubrics grounded in guidelines or validated by experts, adjudicators independent of both the sponsor and the model developer. Then comes the sentence I read twice, because it puts in writing that an LLM judge is admissible if it is independent and qualified: those independence expectations "would still be applicable when the expert adjudicator is itself an LLM." In October 2024 I asked FDA that question in a published pre-submission: can an LLM that is not the device model label a secondary test set? The paper also invites synthetic data and "virtual patient avatars," and names the problems with public benchmarks: contamination, saturation, weak real-world representativeness. Both were in the foundation model FAQ I published in July 2024.
I have already run the exam FDA is describing, just on a different examinee. In July I published a 1,200-question regulatory judgment benchmark that graded more than two dozen frontier AI configurations against an opinionated answer key distilled from 15 years of SaMD consulting, with free-text answers scored by a three-judge AI panel on a two-of-three majority and a strict held-out set. That is FDA's proposed structure one section at a time: an expert-consensus comparator, an independent LLM adjudicator, sequestered items to defeat contamination, and prespecified pass rules. It also produced the lesson the paper's benchmarking section will need. Putting all 22 models to a vote scored 66 percent, worse than the best single model, because the models share their wrong answers, and every model failed the questions with the highest cost of error. A committee of judges is not independence, and raw accuracy hides the tail that maps to FDA's consequences axis. If you build a device benchmark, weight the items by what a wrong answer costs and hold out the ones that decide the expensive calls.
Benchmarking "may not fully establish how a GenAI-enabled device will perform in real clinical use," so FDA adds clinical confirmation, with explicit relief: it "might not require a prospective clinical study in every case." Five approaches, in rising rigor and patient exposure: retrospective evaluation on real patient inputs (synthetic supplements allowed where data is thin), shadow deployment with outputs recorded but hidden from care, standardized patient interactions with trained actors, blinded or unblinded clinician adjudication of real cases, and a prospective study, sometimes an RCT.
For open-ended outputs with no single correct answer, FDA floats the comparator I have been recommending for generated text: "a panel of qualified clinicians whose consensus reflects the applicable standard of care, or... a median clinician in practice." FDA concedes these designs "may not be powered around traditional effectiveness endpoints" and asks how sample sizes should be set (Question 12). That is an invitation to arrive with a validated LLM-judged error-rate endpoint and a human-adjudicated subsample, sized the way we size sensitivity and specificity studies today. Section VI opens with the sentence that changes the economics of this category: "CDRH is considering whether it is appropriate to accept greater premarket uncertainty regarding a GenAI-enabled device's benefit-risk profile through greater reliance on postmarket monitoring."
The monitoring options are periodic re-benchmarking against the premarket thresholds, sample-based review of real-world inputs and outputs by independent clinicians, and drift monitoring. FDA asks whether "machine-based supervisory agents" can run some of it (Question 20). The premarket assessment becomes the baseline: a modified device is "re-benchmarked against the same capabilities," with a PCCP as one mechanism, and third-party foundation model updates get their own question (Q24) because the vendor rather than the sponsor may trigger the change.
On November 21, 2024, at FDA's first Digital Health Advisory Committee meeting on generative AI, I used my five minutes in the open public hearing to propose that trade: a frontier model generating site-specific synthetic test data, automated ground-truthing, and a nightly comparison against device outputs so drift is caught as it happens, offered as the reason FDA could accept a lighter premarket package. In the Q&A I called the PCCP "test-driven development, but you're getting your tests pre-approved." The pre-submission I had published a month earlier asked FDA whether nightly reruns were enough to detect a cloud vendor silently changing the model. I restated the 90/10 premarket-to-postmarket split in February. Although FDA’s paper adopts none of it as policy, it asks the industry whether it should, which is a step in the right direction, inviting a policy change. I continue to stand by my stance that trading premarket burden for postmarket rigor is better for industry and FDA alike. It trades capex to COGS and ensures safety well after the FDA checkpoint is passed.
Section VII floats a voluntary Foundation Model Device Master File: a model developer submits model or system card information, held confidentially by FDA, that sponsors reference with the holder's permission. FDA's footnote lists what it wants: architecture, training data provenance, healthcare failure modes, subgroup benchmark results, guardrails, "update notification commitments," audit log availability. A MAF would not authorize the model for any use; the sponsor still proves the device.
The idea has been in the air. When I ran my first foundation-model pre-submission in 2024, FDA's own minutes suggested the Master File program for a model used as a component of someone else's device, and the day before DHAC I argued that vendors who publish training-data summaries make their customers easier to clear. Question 25 asks the honest follow-up: given "limited incentive to disclose," what would make a voluntary MAF useful? My answer is contractual. Negotiate version pinning, change notification, and disclosure terms with the model vendor now; a MAF gives that contract a regulatory address.
Agentic systems get a definition (they "autonomously plan and execute multi-step tasks, use external tools, or take actions across a sequence of steps"), a note that documentation and outreach agents may not be device functions, and one benchmarking element of their own. A.1 covers human checkpoints before irreversible steps, and injection through retrieved content and tool outputs.
Nothing at the review desk changes today. The route that works is the one UpDoc used for K253281 in December 2025: a narrow, clinician-supervised function, deterministic clinical logic inside a conversational shell, frozen and versioned models, a scoped PCCP, standard 510(k). What the paper adds is a picture of competency-shaped evidence, and building it now costs little:
We are running this mapping for clients in pre-submissions now, because a Pre-Sub that proposes the competency structure before FDA asks for it sets the agenda for the meeting.
Comment. FDA accepts partial responses, so answer the questions that touch your product: for most LLM device sponsors that is Q7 (is the competency approach right), Q10 (benchmark validity), Q11 and Q12 (confirmation rung and sample size), Q18 (the postmarket trade), Q22 (which changes stay inside the QMS), Q24 (third-party model changes), and Q25 (MAF content). The answers FDA collects here become the evidence bar you are graded against later. Innolitics will file a comment on the docket; if you want a real device scenario represented, anonymized, we will fold it in.
When are comments due on FDA's generative AI discussion paper? October 19, 2026, on Regulations.gov under docket FDA-2026-N-7874. FDA accepts partial responses; you do not need to answer all 26 questions.
Is the FDA generative AI discussion paper binding guidance? No. It is a discussion paper from CDRH's Digital Health Center of Excellence. It is not draft or final guidance and explicitly does not propose policy changes or evidence expectations for marketing submissions.
Does FDA regulate LLMs as medical devices? FDA regulates device functions, not models. An LLM-enabled function with a medical intended use is a device. UpDoc's K253281, cleared December 23, 2025, shows a narrow, clinician-supervised LLM function can clear through a 510(k).
What is FDA's competency-based approach? Non-clinical device benchmarking across safety, clinical proficiency, generalizability, and agentic elements, followed by clinical confirmation ranging from retrospective evaluation to a prospective study, with rigor scaled to the device's position on the two-axis risk framework.
What is a Foundation Model Device Master File? A voluntary, confidential filing under FDA's existing Device Master File program in which a foundation model developer supplies model or system card information that device sponsors can reference in their submissions.
Can an agentic AI medical device get cleared in 2026? As a bounded, supervised function with a predicate or a De Novo path, yes. The paper adds an agentic competency element (A.1) and signals heavier scrutiny of multi-step autonomy and tool use.
FDA press release, August 18, 2026: FDA Seeks Public Feedback to Inform Regulatory Approach for Generative AI-Enabled Medical Devices. FDA discussion paper landing page and PDF. Regulations.gov docket FDA-2026-N-7874. DHAC November 2024 executive summary and Day 2 transcript. Patel B, Blumenthal D. JAMA Health Forum 2026, doi:10.1001/jamahealthforum.2025.6947. Bergman A, Wachter RM, Emanuel EJ. JAMA 2026, doi:10.1001/jama.2026.5483. Freyer O et al. Nature Medicine 2025, doi:10.1038/s41591-025-03841-1. UpDoc K253281 510(k) summary. Innolitics: Foundation models and FDA clearance FAQ (Jul 2024), FDA strategy for foundation models pre-sub (Oct 2024), Coffee talk on generative AI and FDA (Nov 19, 2024), DHAC open public hearing comments (Nov 21, 2024), RSNA talk (Dec 2024), Gen AI device to market without burning runway (Feb 2026), First FDA-cleared LLM-enabled agent (Jun 2026).
Editor note: gate this download behind email signup on the site. It expands the article to 16 pages: all 26 discussion questions mapped to sections, who should answer each, my dated positions, the ten-element benchmarking table with a "what to build" column, the clinical confirmation ladder with rung selection, the DHAC 2024 transcript quotes, the 1,200-question benchmark lesson, a builder's readiness checklist, and a full source list.