Specialty › Bayesian

Robust t-test (model-averaged)

The model-averaged robust Bayesian t-test reports inclusion Bayes factors for an effect, for unequal variances and for heavy tails, averaged over an eight-model ensemble.

What is Robust t-test (model-averaged)?

The plain Bayesian t-test fixes two modelling choices you may be unsure about: it assumes normal data and (often) equal variances. The robust version treats those as questions rather than assumptions. It builds 2×2×2 = 8 models by crossing three binary choices — effect present vs. absent (δ ~ Cauchy(0, r) vs. δ = 0), unequal vs. equal variances (a variance-allocation ρ ~ Beta(3/2, 3/2) vs. ρ = 1/2), and a Student-t vs. normal likelihood (t degrees of freedom ν ~ 2 + Exponential(1), prior mean ν = 3) — then averages over all eight using their posterior model probabilities.

Each of the three questions gets an INCLUSION Bayes factor: the posterior odds that the component is present divided by the prior odds (½ vs. ½, so 1). BF_effect answers 'is there a difference in means?', BF for unequal variances answers 'do the groups differ in spread?', and BF for heavy tails answers 'are there outliers / heavier-than-normal tails?'. Because the Student-t likelihood downweights outliers, the effect Bayes factor is robust to a few extreme observations that would distort a normal-likelihood test.

The model-averaged posterior for the standardized effect δ pools the four effect-present models weighted by their posterior probabilities, so the reported effect size and 95% credible interval already account for uncertainty about variance structure and tail thickness.

Prior model probability is spread equally (1/8 each). Restricting the ensemble to the normal / equal-variance pair recovers the standard JZS two-sample Bayes factor exactly, so this is a strict generalization of the plain Bayesian t-test.

When should I use Robust t-test (model-averaged)?

  • Two independent groups where you are unsure whether the data are normal, whether variances are equal, or whether outliers are present.
  • When you want a single analysis that reports evidence for the effect while explicitly hedging over those modelling choices, rather than committing to one t-test variant.
  • Small-to-moderate samples with possible outliers, where a normal-likelihood test is fragile.

What data does it need?

Two numeric columns (independent groups) + effect prior r-scale.

What does it report?

Three inclusion Bayes factors (effect / unequal variances / heavy tails) with posterior inclusion probabilities, a model-averaged posterior median + 95% CrI on the standardized effect δ, and the 8-row model table (relative log marginal likelihood + posterior probability).

What does it assume?

  • Two independent samples; observations independent within each group.
  • The Cauchy effect prior and the Beta / shifted-exponential structural priors are sensible defaults; rerun with a wider r-scale to check prior sensitivity.

How do I interpret the result?

BF_effect > 1 favours a difference in means; BF for unequal variances > 1 favours heteroscedasticity; BF for heavy tails > 1 favours a Student-t (outlier-robust) likelihood. Read each on the Jeffreys scale (3 / 10 / 100).

A heavy-tail BF well above 1 is a signal that a classical t-test may be distorted by outliers — the model-averaged effect here already discounts them.

The per-model table shows which combinations of assumptions the data actually support; a single dominant row means the averaging is effectively picking one model.

See also

References

  • Maier, Bartoš, Quintana, Marsman, Ly & Wagenmakers (2024). Model-averaged Bayesian t tests. Psychonomic Bulletin & Review.
  • Bartoš & Maier (2022). Robust Bayesian t-test.
  • Kass & Raftery (1995). Bayes factors. JASA 90(430):773-795.
  • Rouder, Speckman, Sun, Morey & Iverson (2009). Bayesian t tests for accepting and rejecting the null hypothesis. PBR 16(2):225-237.