How many samples for a method comparison? There is no single number. Sample size depends on the spread of concentrations, the precision you need on the bias, and the decision points that matter. Range matters as much as count.

“How many samples?” is the first question of any method comparison. The tempting answer is a round number like 40 or 100, and it is the least useful one. CLSI EP09-A3 suggests a minimum of around 40 patient samples, which is a sensible floor. But a count alone misses what the samples are for. Samples exist to estimate the bias at the concentrations where decisions are made, tightly enough to trust, and that takes more than a headcount.

Range matters as much as count

Forty samples clustered in the middle of the measuring interval are worth less than forty spread across it. The bias at a low decision point and a high one can differ entirely. A regression can only pin down the relationship across a range it has actually seen. Extrapolating a fitted line beyond the data is guesswork.

The first requirement is spread: samples reaching both ends of the interval and distributed across it, ideally with enough at and around each medical decision point to estimate the bias there precisely. A study can have plenty of samples and still be underpowered where it matters, if none of them sit near the decision level.

Three rows of sample positions along the measuring interval with a marked decision point. Top: samples clustered in the middle. Middle: samples spread evenly across the range. Bottom: samples spread across the range with an extra cluster at the decision point.
Placement matters as much as count. Clustered samples leave the extremes to extrapolation. Spreading them across the interval anchors the fit end to end. Adding density around a decision point tightens the bias estimate exactly where a result changes the decision.

Size on the precision you need

The right way to think about the count is backwards from the answer. The output is the bias at each decision point, with a confidence interval. How tight that interval needs to be is a clinical judgement. The interval must be narrow enough to place the bias confidently inside or outside the allowable limit.

A wider decision matters less and tolerates fewer samples. A bias sitting close to the allowable limit needs a tight interval, and therefore more. Sizing the study on the interval width you require, rather than on a habit or a minimum, is what stops you from finishing a comparison that cannot actually answer the question.

Replicates, and what they buy

Measuring each sample in replicate on both methods reduces the contribution of measurement error to the comparison, sharpening the estimate of the true relationship without collecting more patients. Replication is not a substitute for range or count, though. Replicates of the same clustered samples still leave the extremes unexamined. Where samples are scarce, replication recovers some precision. The choice between more samples and more replicates depends on which is the binding constraint: patient availability, or the measurement noise of the assays.

Downloads

Download the CLSI EP09-A3 method comparison example workbook (.xlsx) — 79 samples across the range with the bias at a decision point and its confidence interval, ready to open in the Analyse-it trial.

Common mistakes

Treating the minimum as the target. Forty is a floor, not a goal. Size the study on the precision your decision needs, which is often more.

Collecting count without range. Samples clustered in the middle cannot estimate bias at the extremes. Spread them across the interval and around the decision points.

Ignoring the decision points when sampling. The bias that matters is at the decision level. Make sure samples sit there, not just somewhere on the range.

Finishing before checking the interval width. A completed study with a decision-point bias interval too wide to act on has not answered the question. Judge sufficiency by the interval, not the count.

Size a method comparison with Analyse-it

Plan the study so the interval lands where you need it; Analyse-it then shows whether it did, inside Excel:

  • Bias at your decision points, each with a confidence interval
  • The interval width, which is what tells you whether the study was large enough to decide
  • The same analysis run on the pilot and on the full study, so the two are comparable

Every feature from all five editions for 15 days. The full method comparison analysis, including the regression fits, is in the Method Validation and Ultimate editions, from US$ 475 a year; the Medical edition covers Bland–Altman agreement and bias at decision points, from US$ 340 a year. Validated against NIST and CLSI reference datasets. See bias at a decision point, or how many samples for a diagnostic accuracy study for the qualitative counterpart.