Guides by field

Clinicians & epidemiologists

Compare treatments, measure risk and rates, follow patients over time, and pool the evidence.

Clinical and population research draw on the same toolbox — the questions differ mainly in scale. Does treatment A beat treatment B, does an exposure raise risk, how long until an event happens, and what does the whole literature say once you pool it? This guide maps those questions onto the right test so you spend your time interpreting results, not hunting for the correct name.

A rule of thumb throughout: match the test to the type of outcome (a number, a yes/no, a count or rate, or a time-to-event) and to the design (independent groups, the same patients measured twice, matched pairs, or person-time follow-up). Keep three roles distinct, because each is handled differently: a confounder distorts an association (adjust or stratify for it), a mediator sits on the causal path (don’t adjust it away if you want the total effect), and an effect modifier changes the effect across strata (report it, don’t pool over it).

Let the study’s reporting standard steer the analysis — CONSORT for randomized trials, STROBE for observational studies, STARD for diagnostic-accuracy studies, TRIPOD for prediction models, and PRISMA for evidence syntheses (all curated by the EQUATOR Network). Two principles from ICH E9 recur in trials: analyze by intention-to-treat, and pre-specify the primary endpoint (and any non-inferiority margin) before looking at the data. Grade the resulting certainty with GRADE.

Comparing an outcome between two groups (CONSORT · ICH E9)

If the outcome is a measurement (blood pressure, HbA1c, a lab value) and each patient appears once, a two-sample t-test compares the group means. If the same patients are measured before and after, use the paired t-test — it is far more powerful because it cancels each patient’s baseline. When the numbers are skewed or ordinal (pain scores, small samples), switch to the rank-based non-parametric versions.

  • — Two independent groups, roughly symmetric measurements.
  • — Before/after on the same patient — the workhorse for within-subject change.
  • — Two groups, skewed or ordinal data (the non-parametric two-sample test).
  • — Paired data that isn’t normal.

Comparing three or more groups (ICH E9)

One-way ANOVA generalizes the t-test to several arms; if you want to adjust for a baseline covariate (e.g. compare final values while controlling for the starting value), ANCOVA is the tool. For measurements repeated over several visits, repeated-measures ANOVA keeps the within-patient structure. The rank-based Kruskal-Wallis is the non-parametric fallback.

  • — Several independent arms; add a post-hoc (Tukey/Dunnett) for pairwise comparisons.
  • — Compare groups while adjusting for a baseline covariate.
  • — The same patients measured across multiple time points.
  • — Non-parametric alternative for skewed outcomes.

Risk, odds & 2×2 measures of association (STROBE)

For a yes/no outcome (cured / not cured, event / no event), a 2×2 table gives you relative risk, odds ratio, risk difference, and number-needed-to-treat in one place — relative risk for cohort designs, odds ratio for case-control. When exposure is a dose with ordered levels, a trend test is more powerful than a plain chi-square. For paired yes/no data (the same patients under two conditions), use McNemar’s test.

  • — RR, OR, risk difference and NNT from a single 2×2 table.
  • — Association between two categorical variables (large samples).
  • — Small cell counts — the exact version of the chi-square test.
  • — Paired yes/no data (e.g. the same patients under two conditions).
  • — The same patients under three or more conditions (McNemar extended).
  • — Trend across ordered exposure levels.

Incidence rates and person-time (STROBE)

When outcomes accrue over person-time, work with incidence rates rather than simple proportions. Compute an exact Poisson confidence interval for a single rate, or compare two rates directly; negative binomial regression handles overdispersed counts where the variance exceeds the mean.

  • — Exact Poisson CI for an incidence rate.
  • — Compare two incidence rates.
  • — Overdispersed count data (variance > mean).

Adjusting for confounding (STROBE · Hernán & Robins)

To adjust for confounders you either stratify or model. Cochran-Mantel-Haenszel pools 2×2 tables across strata; logistic and Poisson regression adjust for several covariates at once and yield adjusted odds or rate ratios. For matched case-control designs, use conditional logistic regression. Propensity-score matching approximates a randomized comparison from observational data.

  • — Stratified 2×2 analysis (Mantel-Haenszel pooled OR).
  • — Adjusted odds ratios with several predictors.
  • — Rate/count outcomes (person-time as offset).
  • — Matched case-control designs.
  • — An outcome with three or more unordered categories.
  • — Repeated measures, or patients clustered within centres.
  • — Balance covariates in observational studies before comparing.

Time-to-event (survival) analysis (CONSORT · STROBE)

When the outcome is “how long until X happens” and some patients haven’t had the event yet (censoring), you need survival methods — the same machinery in a trial or at population scale. Kaplan-Meier draws the survival curves and compares groups with the log-rank test; Cox regression estimates hazard ratios while adjusting for covariates. Restricted mean survival time is a friendlier summary when the curves cross or proportional hazards is doubtful.

  • — Survival curves + log-rank test, with median and per-group summary.
  • — Adjusted hazard ratios from proportional-hazards regression.
  • — Proportional hazards in doubt? Let a covariate’s effect vary with time.
  • — Exposures that change during follow-up (start/stop counting-process data).
  • — Restricted mean survival — robust when hazards aren’t proportional.
  • — Competing risks (e.g. death from another cause before the event).
  • — Accelerated-failure-time fit when you need to extrapolate beyond follow-up.

Diagnostic & screening tests (STARD)

A ROC curve summarizes how well a continuous marker separates disease from no-disease across every possible cutoff, with the AUC as a single headline number. To compare two markers on the same patients, DeLong’s test tells you whether one AUC is genuinely higher. Screening programs also need predictive values, which depend on prevalence: if you already have sensitivity, specificity, and prevalence, the diagnostic calculator turns them into the values a clinician actually cares about.

  • — AUC, the best cutoff, sensitivity/specificity, and the AUC ≠ 0.5 test.
  • — Is marker A’s AUC better than marker B’s? (DeLong, paired).
  • — Sensitivity, specificity, PPV, NPV and Youden’s J as functions of the cutoff.
  • — AUC restricted to the false-positive range you can actually tolerate.
  • — Precision-recall curves — the better view when disease is rare.
  • — Likelihood ratios per result band, not just above/below one cutoff.
  • — Turn sensitivity/specificity/prevalence into PPV, NPV, and likelihood ratios.
  • — Compare two AUCs from published summary statistics alone.

Do two measurement methods agree? (Bland–Altman 1986)

Correlation is the wrong tool for comparing a new device against a reference — two methods can correlate perfectly while being systematically biased. Bland-Altman shows the bias and limits of agreement directly, and Passing-Bablok / Deming regression handle error in both variables.

  • — Bias and 95% limits of agreement between two methods.
  • — Robust method-comparison regression (no normality assumption).
  • — Agreement between two raters on a categorical rating.

Synthesizing the evidence (meta-analysis) (PRISMA · MOOSE)

Once several studies address the same question, pool them with random- or fixed-effect meta-analysis, choosing the input format that matches what the papers report (2×2 counts, means/SDs, single proportions, or a generic effect + SE). Report heterogeneity (I², τ²) and grade the pooled certainty; a robust Bayesian option adjusts for publication bias.

  • — Generic inverse-variance pooling when papers report an effect and its SE.
  • — Pool 2×2 counts into a summary RR/OR (with heterogeneity).
  • — Pool means/SDs across studies.
  • — Pool single-proportion estimates.
  • — Publication-bias-adjusted Bayesian pooling.

Planning the study (ICH E9)

Before collecting data, a power/sample-size calculation tells you how many patients you need to detect a clinically meaningful effect. Pick the calculator that matches your primary endpoint.

  • — Sample size for comparing two means.
  • — Sample size for comparing two proportions/rates.
  • — Events/sample size for a survival endpoint.
  • — Show a new treatment is no worse than the standard by more than Δ.
  • — Equivalence of two means — both one-sided tests inside ±Δ.

References & further reading

Other fields: · · · · · · · ·