Grubbs' test tests whether the single most extreme value in a sample is a statistical outlier under a Normal model.
Grubbs computes how many SDs the most-extreme value lies from the mean (G = |max − mean| / SD or |min − mean| / SD) and compares to a critical value tied to the t-distribution. The test gives a clean yes/no on one suspect value while controlling Type-I error analytically — useful when Tukey's rule has flagged a single observation and you want a formal test before deleting it.
The critical limitation is that Grubbs assumes the rest of the data is Normal. If two or more genuine outliers are present, the SD they inflate masks each other (the so-called masking effect) and Grubbs can fail to flag any of them. For more than one suspected outlier, use ROUT instead.
One numeric column + significance level α (default 0.05).
G statistic (deviations from the mean in SDs), the suspect value, and a p-value or critical-value verdict.
If Grubbs is significant, you have evidence the most extreme value is not part of the assumed Normal population. Document the reason for any deletion (data-entry error, instrument failure, etc.) rather than deleting on the test result alone.
If you suspect ≥ 2 outliers, do not apply Grubbs iteratively (delete the worst, re-test, delete next, …) — the SE is now estimated on the cleaned data and the test is no longer valid. Use ROUT, which is FDR-controlled and handles multiple outliers natively.