The two-sample t-test asks whether two independent groups have equal means — Welch's version (unequal variances) by default, with classical Student's t available.
The two-sample t-test compares two group means. Two flavours: classical Student's assumes equal variances and pools them; Welch's allows unequal variances and Satterthwaite-adjusts the df. Welch is the safer default — it has almost no power loss when variances are equal and protects you when they aren't (Ruxton 2006).
The non-parametric alternative is the Mann-Whitney U / Wilcoxon rank-sum test. The big distinction: Mann-Whitney tests whether one distribution is stochastically larger (P(X > Y) ≠ 0.5), not whether the means differ. If you specifically want to compare means, the t-test is the right tool even on slightly non-normal data — robustness via the CLT kicks in by n ≈ 30 per group.
For paired observations (before/after on the same subject), use the paired t-test instead — pairing dramatically reduces variance and you'd be throwing away that structure with a two-sample test.
Two numeric columns (one per group).
t statistic, df (Welch-adjusted), two-sided p-value, mean difference + 95% CI, three effect sizes (Cohen's d, Hedges' g, Glass's Δ).
Report the mean difference + CI before the p-value — magnitude first, significance second. A 'significant' difference of 0.3 mmHg is rarely clinically meaningful.
Cohen's d ≈ 0.2 small, 0.5 medium, 0.8 large — but these are calibration guidelines, not laws of nature. In a specific field, what counts as a meaningful effect is set by domain knowledge, not by Cohen.