Compare › Parametric

Two samples t-test

The two-sample t-test asks whether two independent groups have equal means — Welch's version (unequal variances) by default, with classical Student's t available.

What is Two samples t-test?

The two-sample t-test compares two group means. Two flavours: classical Student's assumes equal variances and pools them; Welch's allows unequal variances and Satterthwaite-adjusts the df. Welch is the safer default — it has almost no power loss when variances are equal and protects you when they aren't (Ruxton 2006).

The non-parametric alternative is the Mann-Whitney U / Wilcoxon rank-sum test. The big distinction: Mann-Whitney tests whether one distribution is stochastically larger (P(X > Y) ≠ 0.5), not whether the means differ. If you specifically want to compare means, the t-test is the right tool even on slightly non-normal data — robustness via the CLT kicks in by n ≈ 30 per group.

For paired observations (before/after on the same subject), use the paired t-test instead — pairing dramatically reduces variance and you'd be throwing away that structure with a two-sample test.

When should I use Two samples t-test?

  • Two independent groups with a continuous outcome.
  • Welch (default) when group variances might differ — even mildly.
  • Classical Student when you have a strong a-priori reason to assume equal variances (rare in practice).
  • Switch to Mann-Whitney for ordinal data, heavily skewed small samples, or when the interest is in distributional shift rather than mean shift.

What data does it need?

Two numeric columns (one per group).

What does it report?

t statistic, df (Welch-adjusted), two-sided p-value, mean difference + 95% CI, three effect sizes (Cohen's d, Hedges' g, Glass's Δ).

What does it assume?

  • Independent samples within and between groups.
  • Approximately normal within each group (CLT covers n ≥ 30 per group).
  • Equal variances (classical Student only — Welch relaxes this).

Formula

Welch t = (x̄₁ − x̄₂) / √(s₁²/n₁ + s₂²/n₂)
df (Welch) = (s₁²/n₁ + s₂²/n₂)² / [(s₁²/n₁)²/(n₁−1) + (s₂²/n₂)²/(n₂−1)]

How do I interpret the result?

Report the mean difference + CI before the p-value — magnitude first, significance second. A 'significant' difference of 0.3 mmHg is rarely clinically meaningful.

Cohen's d ≈ 0.2 small, 0.5 medium, 0.8 large — but these are calibration guidelines, not laws of nature. In a specific field, what counts as a meaningful effect is set by domain knowledge, not by Cohen.

See also

References

  • Welch (1947). The generalization of Student's problem when several different population variances are involved. Biometrika 34(1/2).
  • Ruxton (2006). The unequal variance t-test is an underused alternative. Behavioral Ecology 17(4).