One-way analysis of variance extends the two-sample t-test to k ≥ 2 independent groups by partitioning total variance into between-group and within-group components.
ANOVA tests the global null that all k group means are equal. The F statistic is the ratio of between-group variance (MS_between) to within-group variance (MS_within); under H₀ it follows F(k−1, N−k). Significant F means at least one group differs from at least one other — but doesn't tell you which.
Three variants are available here for the omnibus test: classical (equal variances), Welch (unequal variances, Satterthwaite df), and Brown-Forsythe F* (also heteroscedastic, slightly different small-sample behaviour). Welch is the right default unless you have strong a-priori reason to assume equal variances — same logic as for the t-test.
Post-hoc pairwise tests follow once the omnibus is significant. Tukey HSD (default) controls FWER for all pairwise comparisons under equal variances. Games-Howell does the same under unequal variances. Dunnett compares everything against one control level. Use Bonferroni / Holm / BH-adjusted pairwise t's when you have only a few pre-specified comparisons.
For ranked / non-normal data, the Kruskal-Wallis test is the rank-based counterpart. For repeated measurements on the same subject, use repeated-measures ANOVA instead.
Numeric response + categorical grouping column.
F, df, p, η², ω². Post-hoc panel with pairwise differences, CIs, adjusted p, significance stars.
η² (eta-squared) = SS_between / SS_total: proportion of variance explained by group. ω² is the same idea with a bias correction; prefer ω² for reporting.
η² ≈ 0.01 small, 0.06 medium, 0.14 large (Cohen's rough guidance). Read these as calibration anchors, not cutoffs.
A significant omnibus + null Tukey HSD is normal when several groups differ from one another by similar small amounts — the omnibus picks up the structure that no single pair captures.