Describe › Outliers

ROUT method

The ROUT method (Motulsky & Brown 2006) identifies outliers from residuals around the median with a Benjamini-Hochberg FDR step-up at a chosen Q, tolerating several outliers at once.

What is ROUT method?

ROUT was designed for non-linear regression residuals from a robust fit, but the same machinery applies to a univariate sample by centring on the median. The key contrast with Grubbs is that ROUT uses the Benjamini-Hochberg step-up to control the false discovery rate across multiple suspected outliers, so masking (where two outliers hide each other from a sequential test) doesn't occur.

RSDR (robust SD) is the 68.27th percentile of the absolute residuals, which equals σ for a standard normal but isn't inflated by outliers. Each observation's residual is divided by RSDR to get a t-like statistic; the sorted p-values are then walked through the BH ladder p₍ₖ₎ ≤ (k/n)·Q, and everything at or below the largest k passing the threshold is flagged.

Q is the FDR — the expected fraction of false flags among all flagged points. Q = 1% gives very conservative flagging (almost no false positives, but real outliers will be missed in heavy-tailed data); Q = 10% flags more freely. Q = 1% is a conservative default and a sensible starting point for clinical / lab data.

When should I use ROUT method?

  • When two or more outliers are suspected — Grubbs and Dixon are unreliable under masking.
  • When the per-outlier α framing isn't what you want; FDR is more interpretable across a sweep.
  • As the screening step before nonlinear regression — ROUT-on-residuals is the original use case.

What data does it need?

One numeric column + Q ∈ {0.001, 0.01, 0.05, 0.1}.

What does it report?

RSDR + list of flagged values + option to remove them in place.

What does it assume?

  • Independent observations.
  • Bulk of the data is approximately normal around the median; ROUT tolerates this assumption better than Grubbs because RSDR is robust to outliers.

Formula

RSDR = 68.27th percentile of |xᵢ − median(x)|
tᵢ = |xᵢ − median(x)| / RSDR, then 2-sided p ∼ t_{n−1}
BH: flag if p₍ₖ₎ ≤ (k / n) · Q for the largest k that passes

How do I interpret the result?

Q is not α: setting Q = 1% means at most 1% of flagged points are expected to be false positives, not that each point has a 1% type-I rate. With Q = 10% and 50 flags you'd expect ~5 to be false alarms.

When ROUT flags more than ~5% of the sample on Q = 1%, the issue is usually distributional (heavy tails, mixture of populations) rather than individual outliers — the right response is a transformation or a robust model, not deletion.

See also

References

  • Motulsky & Brown (2006). Detecting outliers when fitting data with nonlinear regression — a new method based on robust nonlinear regression and the false discovery rate. BMC Bioinformatics 7:123.