Guides by field

Planning a study — power & sample size

Work out how many you need before you collect — for significance, for precision, or to show equivalence.

The most consequential statistics in a study happen before any data exists. An underpowered study wastes the participants it recruited and answers nothing; an oversized one wastes resources and, in a clinical setting, exposes people to risk for no gain. This guide covers sizing a study for whichever question you are actually asking.

There are two ways to size a study, and they answer different questions. Powering for *significance* asks "how many do I need to detect an effect of size Δ, if it is real?" — you supply the smallest effect worth detecting, not the effect you hope to see. Sizing for *precision* asks "how many do I need for my confidence interval to be no wider than W?" — the better frame whenever you are estimating a quantity (a rate, a limit of agreement, a reference limit) rather than testing a hypothesis.

Two conventions keep the result honest: pre-specify the primary endpoint and the analysis before you look at the data (ICH E9), and register the plan where your field expects it. Post-hoc "observed power" computed from your own result adds nothing — it is a restatement of the p-value, not evidence about the design.

Powering the classical comparisons (ICH E9)

Pick the calculator that matches the primary endpoint and design. Each solves for whichever quantity you leave blank — sample size from a target power, or power from a sample size you are stuck with.

  • — Comparing two means (or one against a target).
  • — Comparing two proportions or event rates.
  • — Three or more groups, sized with Cohen's f.
  • — Detecting a correlation different from zero.
  • — Goodness-of-fit or independence, sized with Cohen's w.
  • — Paired binary outcomes (the same subjects twice).

Diagnostic and time-to-event endpoints (STARD · CONSORT)

A diagnostic study is sized on the AUC, and a survival trial on the number of *events* rather than the number of patients — which is why follow-up length trades off against recruitment.

  • — Is a single marker better than chance (AUC ≠ 0.5)?
  • — Comparing two markers measured on the same patients.
  • — Events (and hence patients) for a survival endpoint.

Sizing for precision instead of significance

When the goal is to *estimate* something — a prevalence, a bias, a reference limit — power is the wrong frame. Size the study so the confidence interval is narrow enough to be useful, by naming the half-width you can live with.

  • — CI for a single mean to a target half-width.
  • — CI for a difference between two means.
  • — CI for a single proportion.
  • — CI for a difference between two proportions.
  • — CI for a paired (McNemar) proportion difference.
  • — CI for a Bland-Altman limit of agreement.
  • — CI for a regression-based reference limit.
  • — Duplicate pairs needed to estimate a CV to target precision.

Showing equivalence or non-inferiority

A non-significant difference is not evidence of no difference — absence of evidence is not evidence of absence. To claim two treatments are interchangeable you must pre-specify a margin Δ and show the whole confidence interval falls inside it. Choosing that margin is a clinical judgement, not a statistical one, and it belongs in the protocol.

  • — Two one-sided tests: equivalence of two means within ±Δ.
  • — Equivalence of two proportions within ±δ.
  • — One-sided: the new treatment is not worse by more than Δ.

References & further reading

Other fields: · · · · · · · ·