Describe › Normality

Shapiro–Wilk test

Shapiro–Wilk tests whether a sample comes from a normally distributed population — the default normality screen for samples up to a few thousand observations.

What is Shapiro–Wilk test?

Shapiro–Wilk compares the ordered observations to what you'd expect if the data were normal. The test statistic W is essentially the squared correlation between the sorted data and the expected normal order statistics — W ≈ 1 when the data hug the Q-Q line, W → 0 as the data depart from it.

It is widely considered the most powerful general-purpose normality test for small to moderate samples (n ≤ 5000), beating Kolmogorov–Smirnov, Lilliefors, and Anderson–Darling on a wide range of alternatives. For very large samples (n in the thousands) the test becomes hypersensitive — it will flag tiny, practically meaningless departures from normality, so visual checks via a Q-Q plot become more useful than the p-value.

Use Shapiro–Wilk as the default normality screen before parametric tests (t-tests, ANOVA, linear regression). If the variable fails, consider a transformation (log, sqrt) or switch to the non-parametric counterpart (Mann–Whitney, Kruskal–Wallis).

When should I use Shapiro–Wilk test?

  • Default normality test for n between roughly 3 and 5000.
  • Before any parametric test that assumes normality of the variable itself (one-sample t, paired-difference t) or of the residuals (regression, ANOVA).
  • When you want a single global verdict rather than a tail-vs-centre breakdown.

What data does it need?

One numeric column.

What does it report?

W statistic (closer to 1 = more normal) and a two-sided p-value. Small p (< 0.05) rejects normality.

What does it assume?

  • Independent observations.
  • Continuous data — ties degrade power because they distort the order statistics.

Formula

W = (Σᵢ aᵢ · x₍ᵢ₎)² / Σᵢ (xᵢ − x̄)²

How do I interpret the result?

A small p rejects normality but tells you nothing about why. Pair the test with a Q-Q plot to see whether the departure is skewness, heavy tails, or a few outliers. Each pattern suggests a different fix: log-transform for right skew, robust methods for heavy tails, ROUT / Grubbs for outliers.

When p is large, you have not proved normality — you have failed to reject it. For small samples this happens routinely even when the underlying distribution is clearly non-normal, simply because the test has low power. Treat W ≈ 1 as supportive evidence, never as proof.

See also

References

  • Shapiro & Wilk (1965). An analysis of variance test for normality. Biometrika 52(3/4).