Describe › Outliers

Grubbs' test

Grubbs' test tests whether the single most extreme value in a sample is a statistical outlier under a Normal model.

What is Grubbs' test?

Grubbs computes how many SDs the most-extreme value lies from the mean (G = |max − mean| / SD or |min − mean| / SD) and compares to a critical value tied to the t-distribution. The test gives a clean yes/no on one suspect value while controlling Type-I error analytically — useful when Tukey's rule has flagged a single observation and you want a formal test before deleting it.

The critical limitation is that Grubbs assumes the rest of the data is Normal. If two or more genuine outliers are present, the SD they inflate masks each other (the so-called masking effect) and Grubbs can fail to flag any of them. For more than one suspected outlier, use ROUT instead.

When should I use Grubbs' test?

  • Approximately normal data with exactly one suspected extreme value.
  • Before deleting a single observation, when you want a defensible p-value rather than just a fence rule.

What data does it need?

One numeric column + significance level α (default 0.05).

What does it report?

G statistic (deviations from the mean in SDs), the suspect value, and a p-value or critical-value verdict.

What does it assume?

  • Independent observations.
  • The rest of the data are normally distributed.
  • Exactly one outlier is suspected — masking destroys Grubbs in the presence of multiple outliers.

Formula

G = max(|xᵢ − x̄|) / s

How do I interpret the result?

If Grubbs is significant, you have evidence the most extreme value is not part of the assumed Normal population. Document the reason for any deletion (data-entry error, instrument failure, etc.) rather than deleting on the test result alone.

If you suspect ≥ 2 outliers, do not apply Grubbs iteratively (delete the worst, re-test, delete next, …) — the SE is now estimated on the cleaned data and the test is no longer valid. Use ROUT, which is FDR-controlled and handles multiple outliers natively.

See also

References

  • Grubbs (1969). Procedures for detecting outlying observations in samples. Technometrics 11(1).