Shapiro–Wilk tests whether a sample comes from a normally distributed population — the default normality screen for samples up to a few thousand observations.
Shapiro–Wilk compares the ordered observations to what you'd expect if the data were normal. The test statistic W is essentially the squared correlation between the sorted data and the expected normal order statistics — W ≈ 1 when the data hug the Q-Q line, W → 0 as the data depart from it.
It is widely considered the most powerful general-purpose normality test for small to moderate samples (n ≤ 5000), beating Kolmogorov–Smirnov, Lilliefors, and Anderson–Darling on a wide range of alternatives. For very large samples (n in the thousands) the test becomes hypersensitive — it will flag tiny, practically meaningless departures from normality, so visual checks via a Q-Q plot become more useful than the p-value.
Use Shapiro–Wilk as the default normality screen before parametric tests (t-tests, ANOVA, linear regression). If the variable fails, consider a transformation (log, sqrt) or switch to the non-parametric counterpart (Mann–Whitney, Kruskal–Wallis).
One numeric column.
W statistic (closer to 1 = more normal) and a two-sided p-value. Small p (< 0.05) rejects normality.
A small p rejects normality but tells you nothing about why. Pair the test with a Q-Q plot to see whether the departure is skewness, heavy tails, or a few outliers. Each pattern suggests a different fix: log-transform for right skew, robust methods for heavy tails, ROUT / Grubbs for outliers.
When p is large, you have not proved normality — you have failed to reject it. For small samples this happens routinely even when the underlying distribution is clearly non-normal, simply because the test has low power. Treat W ≈ 1 as supportive evidence, never as proof.