Specialty › Bayesian

t-test (Bayes factor)

The JZS (Jeffreys-Zellner-Siow) Bayes factor weighs the evidence for and against a difference in means, for one-sample, two-sample or paired designs.

What is t-test (Bayes factor)?

The standard t-test gives you 'evidence against H₀'; it can never give you evidence *for* H₀ (a large p doesn't mean H₀ is true). The Bayes factor reverses this: BF₁₀ = P(data | H₁) / P(data | H₀). BF₁₀ > 1 favours H₁, BF₁₀ < 1 favours H₀, and the magnitude is interpretable as the relative likelihood under each hypothesis.

The JZS prior on the effect-size scale δ is Cauchy(0, r), with the r-scale picking the prior width: 'medium' (r = √2/2 ≈ 0.707) is the default; 'wide' (r = 1) and 'ultrawide' (r = √2 ≈ 1.414) shift the prior toward larger effects, making BF₁₀ smaller for given data.

Jeffreys' interpretation scale: BF > 10 strong evidence for H₁; BF 3–10 moderate; BF 1–3 anecdotal; BF < 1/3 anecdotal-for-H₀; BF < 1/10 moderate-for-H₀; BF < 1/30 strong-for-H₀.

When should I use t-test (Bayes factor)?

  • When you want to quantify evidence for the null, not just evidence against it.
  • Replication studies — BF₁₀ ≈ 1 honestly reports 'inconclusive' rather than spuriously rejecting.
  • Sequential / optional-stopping designs where p-values mis-calibrate (BFs are decision-time-invariant).

What data does it need?

One or two numeric columns + paired flag + null mean (one-sample) + prior r-scale.

What does it report?

BF₁₀, BF₀₁ = 1/BF₁₀, posterior median + 95% CrI on Cohen's δ, Jeffreys-scale verbal interpretation.

What does it assume?

  • Independent observations within and (for two-sample) between groups.
  • Normality of the data (or CLT-protected n).
  • Prior is a sensible representation of pre-data beliefs about effect size.

How do I interpret the result?

BF₁₀ is multiplicative evidence. BF₁₀ = 10 means the data are 10× more likely under H₁ than under H₀; that's strong evidence by Jeffreys but not 'proof'.

Prior sensitivity matters: rerun with wide and ultrawide r-scale to check the conclusion isn't an artefact of the default.

See also

References

  • Rouder et al. (2009). Bayesian t tests for accepting and rejecting the null hypothesis. PBR 16(2).