How many samples for a diagnostic accuracy study? Sensitivity and specificity are proportions, and a proportion from a few dozen subjects carries a wide confidence interval. You do not size a diagnostic study for the point estimate — you size it so the interval around that estimate is tight enough to act on.

A sensitivity of 90% sounds settled until you see it came from 18 of 20 diseased subjects. The 95% confidence interval runs down to around 70%. The point estimate is not the finding. The interval is. So the design question is not “how many samples give me 90%?” but “how many samples give me 90% ± something I can live with?”

Size the interval, not the point estimate

The width of a proportion’s interval scales with the square root of p(1 − p) / n. Precision improves only with the square root of sample size. Halving the interval costs roughly four times the subjects. Two consequences follow. Proportions near 0.5 need the most subjects. Proportions near 1 need fewer, which is fortunate, because sensitivity and specificity are usually high.

Three 95% confidence intervals around the same 90% sensitivity, at n equals 20, 70 and 139 diseased subjects: the interval narrows from about plus or minus 15% to about plus or minus 5% as n grows.
Same 90% sensitivity, three sample sizes. The point estimate does not change; the interval — which is what you actually report and defend — is what more subjects buy, and it narrows only with the square root of the count.

Sensitivity and specificity are sized separately

Sizing the two arms separately is the step most plans miss. Sensitivity is estimated only from diseased subjects. Specificity only from non-diseased ones. The two arms are independent sample-size problems. The width of each interval is governed by the count in that arm, not by the total. For a rare condition sampled consecutively, the non-diseased arm fills quickly while the diseased arm fills slowly. The diseased arm, the smaller one, dominates the precision of the number you most want to defend. Sizing a study to compare two tests against each other is a different calculation again, because the two tests are measured on the same subjects.

A quick table to plan from

Approximate number of subjects in each arm for a two-sided 95% interval of the stated half-width, at the expected proportion:

Expected sensitivity (or specificity) ±10% ±7% ±5%
80% 62 125 246
90% 35 71 139
95% 19 38 73

Read it as: to pin a sensitivity you expect near 90% to within ±5%, plan for about 139 diseased subjects. And, separately, enough non-diseased subjects to pin specificity to the width you need. The table gives planning figures from the normal approximation. The interval you finally report should use a method that behaves near the boundary, such as Wilson.

Sizing for AUC and for comparisons

If the endpoint is the area under the ROC curve rather than a single operating point, the sample size depends on three things: the AUC you expect, the width you want, and the allocation between diseased and non-diseased groups. A higher expected AUC needs fewer subjects for the same precision. If the study exists to show one test beats another, you are sizing for the difference in AUC. A paired design, with both tests run on the same subjects, needs far fewer than two independent groups, because pairing removes between-subject variation from the comparison.

Consecutive versus enriched sampling

Enrolling consecutive patients preserves the true prevalence and the natural case mix, which protects you from spectrum bias. But for a rare condition it leaves the diseased arm short. Enriching the diseased group fixes the count but distorts prevalence. Predictive values from an enriched study are meaningless unless you re-anchor them to a real prevalence (or report likelihood ratios instead). Decide which you are doing before enrolment, not after.

Common mistakes

Sizing on the total, not the arms. Two hundred subjects with eight diseased gives a precise specificity and a useless sensitivity. Count the arms separately.

Planning for the estimate, not the interval. “We need to show 90%” is not a sample size. “We need 90% to within ±5%” is.

Enriching, then quoting a predictive value. Predictive values depend on prevalence. Enrichment has thrown the prevalence away. Report likelihood ratios, or re-anchor to a real prevalence.

Sizing without the submission in mind. Where the study supports a regulatory claim, the reporting expectations shape the design as much as the target precision does. See performance studies for an IVD 510(k).

Forgetting the spectrum. Easy cases inflate both sensitivity and specificity. A study sized perfectly on clear-cut subjects still misleads about performance in the ambiguous middle where the test is actually used.

Report the precision you achieved with Analyse-it

Plan the study so the intervals land where you need them; Analyse-it then shows the precision you actually got, inside Excel:

  • Sensitivity, specificity, predictive values and likelihood ratios with Wilson confidence intervals (EP12-A2)
  • Area under the ROC curve with DeLong confidence intervals (EP24-A2)
  • Interval widths, which is where an under-powered study shows itself

Every feature from all five editions for 15 days, with no sign-up and no licence key. Diagnostic accuracy is in the Medical, Method Validation and Ultimate editions, from US$ 340 a year. Validated against NIST and CLSI reference datasets. See how many samples for a method comparison for the quantitative counterpart.