{"slug": "what-fda-reviewers-will-ask-about-your-sensitivity-specificity-sample-size", "title": "What FDA Reviewers Will Ask About Your Sensitivity/Specificity Sample Size", "summary": "The U.S. Food and Drug Administration (FDA) does not mandate a universal sample size for sensitivity and specificity studies, but its final 2025 guidance for AI-enabled device Predetermined Change Control Plans (PCCPs) includes performance-evaluation checklist questions on sample size determination, sensitivity/specificity tradeoffs, reference standard management, and clinical rationale for acceptance criteria. An FDA-authored paper shows that 5/5 correct yields 100% sensitivity but a two-sided 95% score interval lower bound of only 56.6%, and even 30/30 raises it to 88.6%, underscoring the importance of denominator size. The article recommends writing the reviewer answer before collecting the first case and provides a conditional power design calculator to ensure at least 90% chance that both lower confidence bounds clear acceptance floors.", "body_md": "If your protocol says, “We will validate the device on 250 patients,” expect the first statistical question to be **250 what?**\n\nThe FDA does not have a universally applicable regulation stating that a particular sensitivity and specificity study should enroll 100, 250, or 500 patients. The specific guidance for a certain device, special controls, risk associated with the claim, and previous FDA feedback on the matter can provide the number to shoot at. Otherwise, it's up to you to link the claim to an adequately specified pass criterion and then demonstrate that the specified number of positives/negatives is likely to result in the study being passed when the device works as intended.\n\nFDA's [final 2025 guidance for AI-enabled device PCCPs](https://www.fda.gov/media/166704/download) is unusually clear on the matter. Its performance-evaluation checklist includes questions about how sample size was determined, the tradeoff between sensitivity and specificity, management of the reference standard and missing data, and the clinical rationale for the acceptance criteria. This checklist was written for modifications to AI-enabled devices through a PCCP. It does not represent a universal checklist for every initial submission. Nevertheless, it closely resembles the set of questions one might encounter during a performance plan stress test.\n\nSmall, seemingly perfect studies demonstrate the importance of the denominator size. An [FDA-authored paper](https://academic.oup.com/cid/article/52/suppl_4/S305/425319) provides the following example: **5/5 correct gives 100% sensitivity, but the lower end of its two-sided 95% score interval is only 56.6%.** Even 30/30 only raises the lower end of the two-sided 95% score interval to 88.6%.\n\nOur recommendation is simple, “**Write the reviewer answer before collecting the first case”.**\n\nSensitivity and specificity are binomial proportions:\n\n\\[\\widehat{Se}=\\frac{TP}{TP+FN}, \\qquad \\widehat{Sp}=\\frac{TN}{TN+FP}.\\]A point estimate alone generally does not make for a good acceptance criterion. Having 90% sensitivity in 20 cases that have the condition present is different from having 90% sensitivity in 200 cases. [FDA guidelines on diagnostic testing](https://www.fda.gov/regulatory-information/search-fda-guidance-documents/statistical-guidance-reporting-results-studies-evaluating-diagnostic-tests-guidance-industry-and-fda) mention that sensitivity and specificity should be presented along with confidence intervals.\n\nFor a binary standalone performance study, a clean demonstration rule is:\n\n\\[LCL_{95\\%}(Se) \\ge Se_{min} \\quad \\textbf{and} \\quad LCL_{95\\%}(Sp) \\ge Sp_{min}.\\]This is just one example of a good design choice that is not a mandatory FDA requirement. Instead, your medical device could require a precision target, comparison, non-inferiority to a predicate, MRMC analysis, or a custom design for the device. In case of a method comparison claim based on a continuous output procedure, a different design problem arises, which is covered by our Bland–Altman sample size guidelines.\n\nIf you plan to apply a one-sided lower confidence limit, state this clearly. A one-sided 95% lower confidence limit is numerically different from the lower limit of a two-sided 95% confidence interval. Below are a calculator and an example of the latter.\n\nA common quick calculation chooses the number of condition-positive and condition-negative cases to target a confidence-interval half-width \\(d\\):\n\n\\[n_+ \\approx \\frac{z_{1-\\alpha/2}^2 Se(1-Se)}{d_{Se}^2}, \\qquad n_- \\approx \\frac{z_{1-\\alpha/2}^2 Sp(1-Sp)}{d_{Sp}^2}.\\]This is a **precision** design. The question asked here is, \"How precise should the estimate be?\"\n\nThe calculator below uses a **conditional power** design. It asks, “If true performance equals our defensible expectation, what case count gives us at least a 90% chance that both lower confidence bounds clear their acceptance floors?” That is usually closer to the business and regulatory question when your protocol has a pass/fail criterion.\n\nThe defaults are an illustrative AI/ML SaMD scenario : expected sensitivity of 0.90, expected specificity of 0.92, lower-bound acceptance floors of 0.80, and a 90% target chance that both endpoints pass. They are useful starting values from our analysis of [AI/ML SaMD acceptance criteria](https://innolitics.com/articles/ai-samd-acceptance-criteria/), not recommendations for every device.\n\nThe calculator treats the positive and negative strata as fixed and independent. To target a 90% chance that both endpoints pass, it targets \\(\\sqrt{0.90}=94.87\\%\\) conditional power for each endpoint. If independence is not defensible, use a more conservative joint-error allocation or simulate the complete design.\n\nBecause exact-binomial power is saw-toothed, this walkthrough uses a conservative monotone-safe quota, the first count for which every larger count through `n_max=10,000`\n\nalso meets the endpoint target. The pointwise first-passing counts are 167 positive and 107 negative cases, but some immediately larger counts fall below target. This convention is a planning choice, not an FDA requirement.\n\nWith the calculator defaults, the targets will be:\n\nAllowing for 10% loss, the mean-yield enriched estimate is **325 patients**. In a consecutive design with 15% condition prevalence, sensitivity is limiting and the mean-yield enrollment estimate is **1,326 patients**. If enrollment must be fixed, **1,445 patients** gives about a 90% probability of obtaining at least 179 evaluable positive cases under the same simple yield assumptions.\n\nConsider now a budget of 250 patients. With 10% losses, this budget leads to roughly 225 evaluable patients. At 15% prevalence, the expected numbers are 33.75 condition-positive and 191.25 condition-negative patients. Specificity may be well covered, but sensitivity is short by 145.25 positive patients, or about 145.\n\nThese are the four numbers that belong in the protocol:\n\n| Number reviewers need | What it answers | Worked example |\n| Evaluable condition-positive cases | Is sensitivity adequately supported? | 179 |\n| Evaluable condition-negative cases | Is specificity adequately supported? | 113 |\n| Mean-yield enriched estimate | How many cases are collected when both strata are filled directly, including losses? | 325 |\n| Consecutive enrollment estimate | How many patients are expected at the planned prevalence and loss rate? | 1,326 |\n\nDo not label all four numbers “sample size.” Each answers a different question.\n\nThis does not inherently make the study impossible to do. Rather, it means that a choice must be made prior to setting the protocol in stone:\n\nWrong choices include maintaining “N=250” in the protocol and hoping the final confidence interval ends up working.\n\nIdentify estimand, analysis population, operating point, interval procedure, confidence level, and inequality. \"Sensitivity will exceed 0.80\" is not a proper statement. \"The lower limits of the specified two-sided 95% exact binomial CIs for patient-level sensitivity and specificity will each be no less than 0.80\" is checkable.\n\nWhen both sensitivity and specificity must pass, consider the success criteria jointly. Under independence, two success criteria with 90% power each have about 81% joint power.\n\nBoth your expected sensitivity and specificity will determine the sample size just as the acceptance floors do. Make sure they are determined using pilot studies with fixed algorithms, a suitable predicate, literature on a similar target population, special controls, or other valid evidence. Do not use final test results to determine the values.\n\nAcceptance criteria should be clinically justified in light of the intended use, the consequences of false-negative and false-positive results, and other relevant evidence. Because those consequences are often unequal, symmetric acceptance floors are not automatically appropriate.\n\nShow \\(n_+\\) and \\(n_-\\) separately. Then show how total enrollment produces both:\n\n\\[N_{enroll} \\approx \\frac{1}{1-r} \\max\\left(\\frac{n_+}{\\pi},\\frac{n_-}{1-\\pi}\\right),\\]where \\(\\pi\\) is the expected condition-positive fraction in the enrolled study flow and \\(r\\) is the pre-specified non-evaluable loss rate. This is an expectation, not a guarantee. The protocol could alternatively specify a minimum number of evaluable subjects for each stratum and collect data until both numbers are reached.\n\n[FDA guidance for diagnostic tests](https://www.fda.gov/regulatory-information/search-fda-guidance-documents/statistical-guidance-reporting-results-studies-evaluating-diagnostic-tests-guidance-industry-and-fda) warns about spectrum bias, including only easy positive cases and clear negative controls can make both measures overly optimistic. Consider relevant disease severities, imitators, confounders, demographics, sites, instrumentation, acquisition methods, and use conditions.\n\nEnrichment is not inherently bad. FDA’s [2022 CADe guidance](https://www.fda.gov/regulatory-information/search-fda-guidance-documents/clinical-performance-assessment-considerations-computer-assisted-detection-devices-applied-radiology) mentions enrichment explicitly as one possibility for an efficient study design. At the same time, it warns about changing reader performance and potential bias from enrichment. Describe the selection criteria, and do not select cases based on how well your system performs on them.\n\nTo report diagnostic sensitivity and specificity directly, condition status generally must be established using a designated reference standard; it need not be a perfect “gold standard.” When the new test is compared only with a non-reference comparator, including a predicate that is not itself a reference standard FDA recommends **positive percent agreement (PPA)** and **negative percent agreement (NPA)** rather than sensitivity and specificity. Increasing sample size alone will not correct bias caused by reference-standard error or by incorporating the candidate test into the reference-standard determination.\n\nExplain who determined the “truth,” what information the determination was based on, and how disagreements or uncertain reference-standard results were resolved. Was the device output part of that reconciliation?\n\nTen lesions in one patient are not automatically ten independent sensitivity observations. Neither are repeated images, video frames, multiple readers, bilateral organs, or several specimens from the same subject. If the claim is patient-level, the patient count is usually the primary denominator.\n\nFor clustered data, use an analysis and power method that preserves the correlation structure, such as cluster bootstrap, GEE, mixed models, or a design-specific simulation. For reader-aided claims, a standalone binomial calculation does not replace an [MRMC study design](https://innolitics.com/articles/mrmc-standalone-analysis/).\n\nSensitivity and specificity trade off against each other. If you pick the threshold after inspecting the test set, it uses information from the test set and makes the confidence interval too optimistic. Set the threshold and all other factors, such as device version, intended use, pre-processing, scoring algorithm, and acceptance criteria, prior to the pivotal trial.\n\nIf you plan to claim multiple thresholds, results, device versions, endpoints, or populations, pre-specify the testing strategy, address multiplicity as applicable, and identify the smallest relevant denominator.\n\nFDA’s diagnostic-test guidance states that discarding equivocal results will likely bias performance estimates. Pre-specify how device failures, low-quality inputs, ambiguous outputs, protocol deviations, missing reference values, and repeat tests will be handled in the primary analysis. Differentiate between an intention-to-diagnose analysis and a per-protocol analysis, and account for every enrolled patient.\n\nIncreasing the sample size to cover 10% loss is only helpful when 10% seems realistic. It doesn’t mean that you are allowed to drop problematic patients silently.\n\nOverall study power does not guarantee meaningful confidence intervals within subgroups of sex, age, race, disease severity, site, scanner, or acquisition protocol. Important cohorts should support appropriate characterization, but subgroup powering is generally unnecessary unless a subgroup claim or known difference makes it necessary. FDA's [2025 sex-specific guidance](https://www.fda.gov/regulatory-information/search-fda-guidance-documents/evaluation-sex-specific-data-medical-device-clinical-studies-guidance-industry-and-food-and-drug) recommends sex-specific analyses of diagnostic performance while noting that subgroup analysis may not be warranted when overall diagnostic accuracy is very high.\n\nSpecify upfront which subgroup analyses are confirmatory, which are descriptive, and when pooling is appropriate.\n\nFor the worked example, the core protocol language could fit in three short paragraphs:\n\n**Primary endpoints and success criterion.** The primary analysis will provide estimates of patient-level sensitivity and specificity at the locked operating point. Success will mean that the lower bounds of the pre-specified two-sided 95% exact binomial confidence intervals for sensitivity and specificity are each at least 0.80.\n\n**Case targets and collection.** With the assumed true sensitivity of 0.90, true specificity of 0.92, and 90% joint conditional-power target under fixed and independent positive and negative strata, the study needs at least 179 condition-positive and 113 condition-negative evaluable patients. Accounting for 10% non-evaluable loss, the mean-yield enriched estimate is 325 patients. Given 15% condition prevalence in a consecutive cohort, the mean-yield enrollment estimate is 1,326 patients. If enrollment must be fixed, 1,445 patients gives about a 90% probability of filling the 179 positive quota under the same simple yield assumptions. For quota-based accrual, collection continues until both minimum evaluable stratum sizes are attained. A fixed-enrollment protocol instead follows its pre-specified total and accrual-monitoring rules.\n\n**Analysis protections.** The reference-standard procedure, analysis unit, operating point, handling of missing and indeterminate results, and clinically important subgroup analyses will be defined before the test-set evaluation.\n\nThese numbers serve as an illustration. Do not copy and paste them into a protocol if your device does not match the assumptions used here. What is important is the logical sequence: **claim → pass rule → expected performance → positive and negative case targets → total enrollment → analysis protections.**\n\nThis implementation uses the same rule as the web calculator. It returns the first monotone-safe case count whose endpoint power remains at or above target for every larger count through `n_max`\n\n.\n\nRun the code blocks in order in one Python session or notebook. The displayed outputs were independently reproduced with Python 3.9.6, NumPy 2.0.2, SciPy 1.13.1, and Matplotlib 3.9.4.\n\n``` python\nfrom dataclasses import dataclass\nfrom math import ceil, sqrt\n\nimport matplotlib.pyplot as plt\nimport numpy as np\nfrom scipy.stats import beta, binom\n\nCONFIDENCE = 0.95\nALPHA_TAIL = (1 - CONFIDENCE) / 2\n\ndef exact_lower_bound(successes, cases):\n    \"\"\"Lower end of a two-sided Clopper-Pearson interval.\"\"\"\n    if successes == 0:\n        return 0.0\n    return beta.ppf(ALPHA_TAIL, successes, cases - successes + 1)\n\ndef required_successes(cases, minimum):\n    \"\"\"Smallest success count whose exact lower bound clears minimum.\"\"\"\n    return int(binom.isf(ALPHA_TAIL, cases, minimum)) + 1\n\nsuccesses = required_successes(cases=30, minimum=0.80)\npower = binom.sf(successes - 1, 30, 0.90)\n\nprint(f\"30/30 exact lower bound: {exact_lower_bound(30, 30):.3f}\")\nprint(f\"Required to clear 0.80: {successes}/30\")\nprint(f\"Power when true sensitivity is 0.90: {power:.1%}\")\n```\n\nOutput:\n\n```\n30/30 exact lower bound: 0.884\nRequired to clear 0.80: 29/30\nPower when true sensitivity is 0.90: 18.4%\n```\n\nThis shows why 30 positive cases are not automatically enough.\n\n```\n@dataclass(frozen=True)\nclass EndpointPlan:\n    cases: int\n    required_successes: int\n    power_at_cases: float\n    lower_ci_at_successes: float\n\ndef exact_endpoint_plan(\n    expected,\n    minimum,\n    endpoint_power,\n    confidence=0.95,\n    n_max=10_000,\n):\n    \"\"\"Return the first count whose power stays at target through n_max.\"\"\"\n    if not 0 < minimum < expected < 1:\n        raise ValueError(\"Require 0 < minimum < expected < 1\")\n    if not 0 < endpoint_power < 1:\n        raise ValueError(\"Require 0 < endpoint_power < 1\")\n    if not 0 < confidence < 1:\n        raise ValueError(\"Require 0 < confidence < 1\")\n\n    alpha_tail = (1 - confidence) / 2\n    case_counts = np.arange(1, n_max + 1)\n    successes = binom.isf(alpha_tail, case_counts, minimum).astype(int) + 1\n    power = binom.sf(successes - 1, case_counts, expected)\n\n    # Exact-binomial power is saw-toothed. Require every later count\n    # through n_max to meet the endpoint target.\n    tail_min_power = np.minimum.accumulate(power[::-1])[::-1]\n    candidates = np.flatnonzero(tail_min_power >= endpoint_power)\n    if len(candidates) == 0:\n        raise ValueError(f\"No monotone-safe solution through n={n_max}\")\n\n    i = int(candidates[0])\n    n = int(case_counts[i])\n    k = int(successes[i])\n    return EndpointPlan(\n        cases=n,\n        required_successes=k,\n        power_at_cases=float(power[i]),\n        lower_ci_at_successes=float(beta.ppf(alpha_tail, k, n - k + 1)),\n    )\n\njoint_target = 0.90  # Illustrative planning choice, not an FDA rule\nendpoint_target = sqrt(joint_target)\n\nsensitivity = exact_endpoint_plan(\n    expected=0.90,\n    minimum=0.80,\n    endpoint_power=endpoint_target,\n)\nprint(sensitivity)\n```\n\nOutput:\n\n```\nEndpointPlan(cases=179, required_successes=154,\n             power_at_cases=0.9659, lower_ci_at_successes=0.8008)\n```\n\nThe sensitivity target is **179 evaluable condition-positive cases**, with at least 154 true positives.\n\n```\nspecificity = exact_endpoint_plan(\n    expected=0.92,\n    minimum=0.80,\n    endpoint_power=endpoint_target,\n)\njoint_power = sensitivity.power_at_cases * specificity.power_at_cases\n\nprint(specificity)\nprint(f\"Joint power at returned counts: {joint_power:.1%}\")\n```\n\nOutput:\n\n```\nEndpointPlan(cases=113, required_successes=99,\n             power_at_cases=0.9639, lower_ci_at_successes=0.8009)\nJoint power at returned counts: 93.1%\n```\n\nThe specificity target is **113 evaluable condition-negative cases**, with at least 99 true negatives. Under fixed and independent strata, the probability that both endpoints pass at the returned counts is 93.1%.\n\n``` python\ndef endpoint_power_curve(expected, minimum, n_max=250):\n    n = np.arange(1, n_max + 1)\n    k = binom.isf(ALPHA_TAIL, n, minimum).astype(int) + 1\n    return n, binom.sf(k - 1, n, expected)\n\nn_se, power_se = endpoint_power_curve(0.90, 0.80)\nn_sp, power_sp = endpoint_power_curve(0.92, 0.80)\n\nplt.figure(figsize=(10, 6))\nplt.plot(n_se, 100 * power_se, color=\"#E85036\",\n         label=\"Sensitivity: expected 0.90\")\nplt.plot(n_sp, 100 * power_sp, color=\"#2D3F86\",\n         label=\"Specificity: expected 0.92\")\nplt.scatter(179, 100 * sensitivity.power_at_cases, color=\"#E85036\")\nplt.scatter(113, 100 * specificity.power_at_cases, color=\"#2D3F86\")\nplt.axhline(100 * endpoint_target, color=\"#059669\", linestyle=\"--\",\n            label=\"94.87% power per endpoint\")\nplt.xlim(20, 250)\nplt.ylim(0, 101)\nplt.xlabel(\"Evaluable cases in the endpoint stratum\")\nplt.ylabel(\"Probability the exact lower bound passes (%)\")\nplt.title(\"Positive and negative case targets are calculated separately\")\nplt.grid(axis=\"y\", alpha=0.25)\nplt.legend()\nplt.tight_layout()\nplt.show()\nprevalence = 0.15\nloss_positive = 0.10\nloss_negative = 0.10\n\npositive_to_collect = ceil(\n    sensitivity.cases / (1 - loss_positive)\n)\nnegative_to_collect = ceil(\n    specificity.cases / (1 - loss_negative)\n)\nenriched_mean_total = positive_to_collect + negative_to_collect\n\nexpected_consecutive = ceil(max(\n    sensitivity.cases / (prevalence * (1 - loss_positive)),\n    specificity.cases / ((1 - prevalence) * (1 - loss_negative)),\n))\n\n# If enrollment must be fixed, power the accrual of the limiting stratum too.\npositive_yield = prevalence * (1 - loss_positive)\nfixed_enrollment_90 = next(\n    n for n in range(expected_consecutive, 10_000)\n    if binom.sf(sensitivity.cases - 1, n, positive_yield) >= 0.90\n)\n\nprint(f\"Mean-yield enriched estimate: {enriched_mean_total}\")\nprint(f\"Expected consecutive enrollment: {expected_consecutive}\")\nprint(f\"90% probability of filling the positive quota: {fixed_enrollment_90}\")\n```\n\nOutput:\n\n```\nMean-yield enriched estimate: 325\nExpected consecutive enrollment: 1326\n90% probability of filling the positive quota: 1445\n```\n\nThe values **325** and **1,326** are mean-yield estimates, not guaranteed fixed sample sizes. Collection should continue until both evaluable quotas are met. If enrollment must be fixed, **1,445** gives about a 90% probability of obtaining at least 179 evaluable positive cases under the same simple yield assumptions.\n\n```\nprevalence_grid = np.linspace(0.05, 0.95, 181)\npositive_driven = sensitivity.cases / (\n    prevalence_grid * (1 - loss_positive)\n)\nnegative_driven = specificity.cases / (\n    (1 - prevalence_grid) * (1 - loss_negative)\n)\n\nplt.figure(figsize=(10, 6))\nplt.plot(100 * prevalence_grid, positive_driven, color=\"#E85036\",\n         label=\"Enrollment needed for 179 positive cases\")\nplt.plot(100 * prevalence_grid, negative_driven, color=\"#2D3F86\",\n         label=\"Enrollment needed for 113 negative cases\")\nplt.plot(\n    100 * prevalence_grid,\n    np.maximum(positive_driven, negative_driven),\n    color=\"#059669\",\n    linewidth=3,\n    label=\"Consecutive enrollment estimate\",\n)\nplt.axvline(15, color=\"#737373\", linestyle=\":\")\nplt.scatter(15, expected_consecutive, color=\"#059669\")\nplt.annotate(\"15% prevalence → 1,326 enrolled\",\n             (15, expected_consecutive), xytext=(25, 15),\n             textcoords=\"offset points\", color=\"#059669\")\nplt.xlim(5, 95)\nplt.ylim(0, 4200)\nplt.xlabel(\"Condition prevalence in consecutive enrollment (%)\")\nplt.ylabel(\"Expected total enrollment\")\nplt.title(\"Prevalence changes enrollment, not the endpoint case targets\")\nplt.grid(axis=\"y\", alpha=0.25)\nplt.legend()\nplt.tight_layout()\nplt.show()\nbudget = 250\nretention = 0.90\nexpected_positive = budget * retention * prevalence\nexpected_negative = budget * retention * (1 - prevalence)\n\nprint(f\"Expected evaluable positive cases: {expected_positive:.2f}\")\nprint(f\"Expected evaluable negative cases: {expected_negative:.2f}\")\nprint(f\"Expected positive-case shortfall: \"\n      f\"{sensitivity.cases - expected_positive:.2f}\")\n```\n\nOutput:\n\n```\nExpected evaluable positive cases: 33.75\nExpected evaluable negative cases: 191.25\nExpected positive-case shortfall: 145.25\n```\n\nThe expected condition-positive count is 145.25 below the 179 case target; because 33.75 is an expected count, it does not imply a guaranteed integer shortfall.\n\nThe 250 subject budget calculation produces 33.75 expected evaluable condition-positive subjects and a 145.25 case expected shortfall relative to the 179-case sensitivity quota. The 179/113 quotas and 325/1,326 planning values come from Steps 2–4, not from this budget check.\n\nThe table below applies to **either** endpoint where each cell is an evaluable condition-positive quota for sensitivity or an evaluable condition-negative quota for specificity, not total enrollment. It uses the lower endpoint of a two-sided 95% Clopper–Pearson interval and 94.87% endpoint design power under assumed true performance, corresponding to a 90% joint target for fixed and independent strata. Values are monotone-safe through `n_max=10,000`\n\n. Here em dash means expected performance does not exceed the floor, so no finite solution exists under this design.\n\n| Expected performance | Minimum 0.70 | Minimum 0.75 | Minimum 0.80 | Minimum 0.85 | Minimum 0.90 |\n| 0.80 | 256 | 929 | — | — | — |\n| 0.85 | 105 | 220 | 776 | — | — |\n| 0.90 | 53 | 89 | 179 |\n595 | — |\n| 0.95 | 31 | 44 | 69 | 127 | 387 |\n| 0.98 | 22 | 27 | 34 | 56 | 114 |\n\nThe following code produces the chart:\n\n```\nfloors = [0.70, 0.75, 0.80, 0.85, 0.90]\ncolors = [\"#059669\", \"#2D3F86\", \"#E85036\", \"#7A5195\", \"#C58A00\"]\n\nplt.figure(figsize=(10, 6))\nfor floor, color in zip(floors, colors):\n    expected_grid = np.arange(floor + 0.01, 0.991, 0.005)\n    required_n = []\n    for expected in expected_grid:\n        try:\n            plan = exact_endpoint_plan(expected, floor, endpoint_target)\n            required_n.append(plan.cases)\n        except ValueError:\n            required_n.append(np.nan)\n\n    plt.plot(\n        100 * expected_grid,\n        required_n,\n        color=color,\n        label=f\"Minimum lower bound {floor:.2f}\",\n    )\n\nplt.scatter(90, 179, color=\"#E85036\", s=60, zorder=5)\nplt.annotate(\n    \"Worked example 0.90 → 179 cases\",\n    xy=(90, 179),\n    xytext=(87.4, 430),\n    textcoords=\"data\",\n    ha=\"center\",\n    va=\"center\",\n    arrowprops={\n        \"arrowstyle\": \"->\",\n        \"color\": \"#E85036\",\n        \"linewidth\": 1.4,\n        \"connectionstyle\": \"arc3,rad=-0.18\",\n    },\n    bbox={\n        \"boxstyle\": \"round,pad=0.35\",\n        \"facecolor\": \"white\",\n        \"edgecolor\": \"#E85036\",\n        \"alpha\": 0.97,\n    },\n    color=\"#E85036\",\n    fontsize=11,\n    zorder=6,\n)\nplt.yscale(\"log\")\nplt.xlim(71, 99)\nplt.ylim(10, 10_000)\nplt.xlabel(\"Expected sensitivity or specificity (%)\")\nplt.ylabel(\"Required evaluable cases (log scale)\")\nplt.title(\"Required sample size rises sharply near the acceptance floor\")\nplt.grid(axis=\"y\", which=\"both\", alpha=0.25)\nplt.legend()\nplt.tight_layout()\nplt.show()\n```\n\nWhat matters is the “cliff.” When expected performance exceeds the acceptance floor by only five percentage points, several hundred cases may be needed. Better performance reduces the required case count, but it should not be an optimistic assumption inserted merely to fit the budget.\n\nThere is no single total. Calculate the required condition-positive cases for sensitivity and condition-negative cases for specificity separately. Then translate those targets into enrollment using prevalence, enrichment, and the expected non-evaluable rate. The larger translated requirement drives the consecutive-study total.\n\nNo universal FDA rule sets 250 as the minimum for every sensitivity/specificity study. Device-specific guidance, a special control, a predicate, a risk profile, or prior FDA feedback may support a particular count. FDA will care whether the positive and negative denominators, confidence intervals, population, and analysis support your exact claim.\n\nPrevalence does not change the binomial positive-case count needed for sensitivity or the negative-case count needed for specificity. It changes how many consecutive patients you must enroll to obtain those counts. In an enriched design, report the fixed stratum counts and explain how the sampling affects representativeness and prevalence-dependent metrics such as PPV and NPV.\n\nBoth appear in FDA's 2007 diagnostic-test guidance: its worked example reports score intervals and points to exact Clopper–Pearson intervals as an alternative. Exact intervals are conservative and behave well near 0 or 1; score intervals can be shorter. Choose the method appropriate for your device and pre-specify it. Do not size with one method and switch after seeing results.\n\nNot as independent observations unless the claim and statistical model justify it. Repeated images, lesions, organs, specimens, and readers are clustered within patients. A naive binomial calculation will make the confidence interval too narrow. Define the claim-level unit and use a clustered analysis or simulation when observations are correlated.\n\nNot automatically. With 30 condition-positive cases, a two-sided 95% exact interval needs at least 29 true positives for its lower bound to clear 0.80. If true sensitivity is 0.90, the chance of observing at least 29/30 is only about 18.4%. A small study can pass, but it is unlikely to pass even when the device performs at its expected rate. Size from the pre-specified rule and desired conditional power, not a customary case count.\n\nThe sample-size section of a reviewer-ready protocol should fit on one page: the sensitivity and specificity claims, their separate lower-bound criteria, the expected performance assumptions and evidence, exact positive and negative case targets, total-enrollment translation, loss assumptions, reference standard, independent unit, subgroup plan, and missing-result analysis.\n\nDo that before data collection and the sample size becomes an engineering decision you can defend. If you start with “N=250” and work backward after the study is locked, you may discover too late that you powered specificity while leaving sensitivity to chance.\n\nTo get a second opinion before Pre-Sub or pivotal data collection, [contact Innolitics](https://innolitics.com/contact/). We review the entire statistical plan together with the intended use, reference standard, data flow, and submission approach, since FDA will do the same.", "url": "https://wpnews.pro/news/what-fda-reviewers-will-ask-about-your-sensitivity-specificity-sample-size", "canonical_source": "https://innolitics.com/articles/what-fda-reviewers-will-ask-about-your-sensitivity-specificity-sample-size/", "published_at": "2026-08-11 05:00:00+00:00", "updated_at": "2026-08-16 05:40:51.692280+00:00", "lang": "en", "topics": ["ai-policy", "ai-products"], "entities": ["U.S. Food and Drug Administration", "FDA"], "alternates": {"html": "https://wpnews.pro/news/what-fda-reviewers-will-ask-about-your-sensitivity-specificity-sample-size", "markdown": "https://wpnews.pro/news/what-fda-reviewers-will-ask-about-your-sensitivity-specificity-sample-size.md", "text": "https://wpnews.pro/news/what-fda-reviewers-will-ask-about-your-sensitivity-specificity-sample-size.txt", "jsonld": "https://wpnews.pro/news/what-fda-reviewers-will-ask-about-your-sensitivity-specificity-sample-size.jsonld"}}