Describe › Normality

Kolmogorov–Smirnov test

The Lilliefors-corrected Kolmogorov–Smirnov test checks normality by comparing the empirical CDF against the normal CDF, with the mean and SD estimated from the data.

What is Kolmogorov–Smirnov test?

The classical Kolmogorov–Smirnov test compares an empirical CDF against a fully specified distribution. When you estimate the parameters (μ, σ) from the same data — as you always do in practice for normality testing — the K–S p-values are too large. The Lilliefors correction adjusts the critical values for this in-sample parameter estimation, giving a properly calibrated test.

K–S statistics are based on the maximum vertical distance between the empirical and theoretical CDFs, so the test is most sensitive in the middle of the distribution and relatively insensitive to tail behaviour. That makes it weaker than Shapiro–Wilk for catching heavy tails, but useful when you specifically care about the centre of the distribution (e.g. calibration of a predicted-vs-observed plot).

When should I use Kolmogorov–Smirnov test?

  • Alternative to Shapiro–Wilk for very large samples (n ≫ 5000) where Shapiro–Wilk is unavailable or hypersensitive.
  • When you want a CDF-based check that's sensitive in the centre rather than in the tails.

What data does it need?

One numeric column.

What does it report?

D statistic (maximum CDF gap) and a Lilliefors p-value. Small p (< 0.05) rejects normality.

What does it assume?

  • Independent observations.
  • Continuous data.

Formula

D = supₓ |F̂(x) − Φ((x − x̄) / s)|

How do I interpret the result?

Less powerful than Shapiro–Wilk in most settings, especially against heavy-tailed alternatives. Reach for K–S only when n is huge or you specifically want the CDF-centre framing — Shapiro–Wilk is the better default for clinical-style data sets.

See also

References

  • Lilliefors (1967). On the Kolmogorov–Smirnov test for normality with mean and variance unknown. JASA 62(318).