The model-averaged robust Bayesian t-test reports inclusion Bayes factors for an effect, for unequal variances and for heavy tails, averaged over an eight-model ensemble.
The plain Bayesian t-test fixes two modelling choices you may be unsure about: it assumes normal data and (often) equal variances. The robust version treats those as questions rather than assumptions. It builds 2×2×2 = 8 models by crossing three binary choices — effect present vs. absent (δ ~ Cauchy(0, r) vs. δ = 0), unequal vs. equal variances (a variance-allocation ρ ~ Beta(3/2, 3/2) vs. ρ = 1/2), and a Student-t vs. normal likelihood (t degrees of freedom ν ~ 2 + Exponential(1), prior mean ν = 3) — then averages over all eight using their posterior model probabilities.
Each of the three questions gets an INCLUSION Bayes factor: the posterior odds that the component is present divided by the prior odds (½ vs. ½, so 1). BF_effect answers 'is there a difference in means?', BF for unequal variances answers 'do the groups differ in spread?', and BF for heavy tails answers 'are there outliers / heavier-than-normal tails?'. Because the Student-t likelihood downweights outliers, the effect Bayes factor is robust to a few extreme observations that would distort a normal-likelihood test.
The model-averaged posterior for the standardized effect δ pools the four effect-present models weighted by their posterior probabilities, so the reported effect size and 95% credible interval already account for uncertainty about variance structure and tail thickness.
Prior model probability is spread equally (1/8 each). Restricting the ensemble to the normal / equal-variance pair recovers the standard JZS two-sample Bayes factor exactly, so this is a strict generalization of the plain Bayesian t-test.
Two numeric columns (independent groups) + effect prior r-scale.
Three inclusion Bayes factors (effect / unequal variances / heavy tails) with posterior inclusion probabilities, a model-averaged posterior median + 95% CrI on the standardized effect δ, and the 8-row model table (relative log marginal likelihood + posterior probability).
BF_effect > 1 favours a difference in means; BF for unequal variances > 1 favours heteroscedasticity; BF for heavy tails > 1 favours a Student-t (outlier-robust) likelihood. Read each on the Jeffreys scale (3 / 10 / 100).
A heavy-tail BF well above 1 is a signal that a classical t-test may be distorted by outliers — the model-averaged effect here already discounts them.
The per-model table shows which combinations of assumptions the data actually support; a single dominant row means the averaging is effectively picking one model.