How many samples do you need to compare two groups? Sample size is not a number you look up. It is the consequence of four quantities that trade against one another — and of one decision nobody wants to make, which is how small a difference you would still care about.

“How many samples do I need?” has no single answer. The number depends on a question only you can settle: how small a difference would still matter enough that you would want to detect it? Fix that, and the arithmetic follows. Refuse to fix it, and no calculation is possible. That is why so many studies are sized by convention, budget, or what fitted on the plate.

Four quantities, one equation

Sample size, the difference you want to detect, the variability of the measurement, and the two error rates are all bound together. Fix any four and the fifth is determined.

The difference worth detecting. Not the difference you hope to find or expect to find, but the smallest one that would change a decision. Setting that difference is a scientific judgement, not a statistical one, and it is the input that most often goes unexamined.

The variability. The standard deviation of the measurement within a group, taken from pilot data, historical records, or a published precision study. Everything downstream depends on it. It’s usually the least certain of the inputs.

The significance level α, the false-positive rate, conventionally 5%.

The power, the chance of detecting the difference if it is really there. Conventionally 80%, which means accepting a one-in-five chance of missing a real effect you have decided you care about. For high-stakes work, 90% is a better choice.

The difference and the variability enter only as their ratio, the standardised effect size: difference divided by standard deviation. That’s convenient. One table covers every measurement scale.

A table to plan from

Approximate subjects per group for a two-sided comparison of two independent means at α = 0.05:

Standardised effect (difference ÷ SD) 80% power 90% power
0.2 — a fifth of a standard deviation 394 527
0.5 — half a standard deviation 64 86
0.8 26 34
1.0 — a full standard deviation 17 23

To use it: if you want to detect a difference of 4 units and the within-group SD is 8 units, the standardised effect is 0.5. Plan for about 64 per group at 80% power. The table gives planning figures only. Add a margin for dropouts, and treat the numbers as approximate when the effect is small.

The square-root rule, and why precision is expensive

The line in that table that matters most is its shape. Halving the effect you want to detect multiplies the required sample by roughly four. Going from 0.5 to 0.2, still not a small difference in most contexts, takes you from 64 per group to nearly 400.

A curve of required sample size per group against standardised effect size, falling steeply from about 400 subjects at an effect of 0.2 to 64 at 0.5, 26 at 0.8 and 17 at 1.0. Two marked points show that halving the effect from 0.5 to 0.25 raises the requirement roughly fourfold.
Required sample size against the effect you want to detect, at 80% power. The curve is a reciprocal square: precision costs quadratically, which is why studies aimed at small effects become large so suddenly.

Two practical consequences follow. Reducing the measurement’s own variability buys the same statistical benefit as increasing the sample, often more cheaply. Better technique, replicate measurements averaged, or a more precise method all count. And a study too small to detect the difference you care about will not simply be inconclusive: it will produce a non-significant result that gets read as evidence of no difference.

Pairing changes everything

For a paired design, the variability that matters is not the variation between subjects but the variation in the within-subject difference, which is usually far smaller. The same table applies, with the standardised effect computed against the SD of the differences.

When subjects differ substantially but respond consistently, pairing can reduce the required sample several-fold. It is the single most effective design change available, and it costs nothing but the ability to measure each unit twice. See Student’s, Welch’s or paired for what pairing does to the analysis.

Two other design points. Equal group sizes are the most efficient allocation for a fixed total. The loss from moderate imbalance is small, but severe imbalance is governed by the smaller group. 180 versus 20 behaves much more like 40 in total than like 200. If the outcome is a proportion rather than a mean, the arithmetic differs. Proportions near 0.5 need the most subjects, and each arm is sized separately, as set out in how many samples for a diagnostic accuracy study.

Post-hoc power is not worth computing

After a non-significant result, it is tempting to compute the power the study had. Do not bother. Power computed from the observed effect is a deterministic function of the p-value, so it carries no information the p-value did not already give you. Such a figure will always be low when the result was non-significant, which makes reporting it circular.

The informative thing to report is the confidence interval on the difference. An interval from −0.5 to +0.7 units, against a difference of 3 units that would have mattered, is genuine evidence that any real effect is small. An interval from −8 to +15 says the study could not tell. That is a quite different statement, and it is the one you should report. Confidence intervals and p-values takes this further.

Prospective power calculations, on the other hand, are worth doing properly and worth writing down before the data exist. The alternative is choosing the analysis after seeing the result.

More than two groups, and multiple endpoints

With three or more groups, size the study on the specific comparison that matters most rather than on the overall F-test. Remember that the multiple comparison procedure widens intervals. A design powered for a single unadjusted comparison will be underpowered once the adjustment is applied. Narrowing the family of comparisons is, once again, the cheapest way to recover it.

The same applies to multiple endpoints. A study with six outcome measures, each tested at 5%, has a much higher chance of a spurious finding somewhere. Declaring one primary endpoint in advance is the usual and best answer.

Common mistakes

Sizing on the effect you expect rather than the smallest one that matters. The first is optimism; the second is the design input.

Computing post-hoc power. It restates the p-value. Report the confidence interval instead.

Using an SD from a source that does not resemble your study. Everything scales with it. A published SD from a different population or method can be badly wrong.

Reading non-significant from a small study as no difference. Absence of evidence, and a wide interval that says so plainly.

Sizing for an unadjusted comparison, then adjusting. The multiple comparison procedure consumes the power you planned.

Forgetting attrition. The sample size is what you need at analysis, not at enrolment.

Report the estimate with Analyse-it

Plan the study so the interval lands where you need it; Analyse-it then reports the estimate that shows whether it did, inside Excel:

  • The mean difference with a t-based or Welch-Satterthwaite confidence interval
  • Cohen’s d and Hedges’ g with non-central t intervals
  • The Hodges-Lehmann location shift for rank-based analyses, in Compare Groups and Compare Pairs

Every feature from all five editions for 15 days, with no sign-up and no licence key. Hypothesis testing is part of the Standard edition, so it is in every Analyse-it edition, from US$ 155 a year. Validated against NIST Standard Reference Datasets. Precision studies to establish the SD your calculation depends on are covered by Method Validation.