The Mann-Whitney U (Wilcoxon rank-sum) test compares two independent groups using ranks rather than means — the non-parametric counterpart to the two-sample t-test.
Mann-Whitney pools both samples, ranks all observations, then asks whether the rank sums differ between groups. Under H₀ the two distributions are identical; under H₁ one group's values tend to be stochastically larger than the other's (P(X > Y) ≠ 0.5).
Common confusion: this is not a 'non-parametric test of equal medians'. It tests for stochastic dominance / location shift only when the two distributions are otherwise the same shape. If shapes differ (e.g. unequal variances), a significant result can come from any distributional difference, not necessarily medians.
Pros vs t-test: no normality assumption, robust to outliers. Cons: lower power than the t-test on normal data, doesn't yield a mean difference (the Hodges-Lehmann estimator we report is the median of all pairwise differences and is the closest non-parametric analog).
Two numeric columns (one per group).
U / W statistic, p-value, Hodges-Lehmann location shift + CI, rank-biserial r (≡ Cliff's δ).
Rank-biserial r runs from −1 to +1: r = 0.5 means group A wins 75 % of pairwise comparisons.
On large samples (n > 50 per group) Mann-Whitney is nearly as efficient as the t-test for normal data; on small samples the t-test wins by ~5 % efficiency. The cost of insurance against non-normality is small.