Experiments, surveys, latent constructs, reliability, and Bayesian alternatives.
Social-science data is a mix of designed experiments and messy surveys, often measuring things you can’t observe directly (attitudes, ability, satisfaction). This guide covers the group-comparison and regression staples plus the latent-variable and reliability tools that psychometrics depends on.
APA’s reporting standards (JARS) ask for effect sizes and confidence intervals alongside p-values, and pre-registration is increasingly expected to guard against p-hacking. Two rules of thumb travel with the methods below: report an effect size (Cohen’s d, η², or an odds ratio) with every test, and treat reliability ≥ 0.70 (Cronbach’s α or McDonald’s ω) as the usual floor for a usable scale.
Comparing conditions (APA JARS)
For experimental designs, the ANOVA family covers between-subject factors, within-subject (repeated) factors, and factorial designs with interactions. Non-parametric versions cover ordinal outcomes like Likert responses.
▸ — Three crossed factors, keeping only the interactions you want.
▸ — Several correlated outcomes tested jointly, before any follow-up ANOVAs.
▸ — Pupils within classes, or sub-samples within subjects — nested, not crossed.
▸ — The two-group case of the same nesting problem.
▸ — Trimmed means: keeps power when the tails are heavy or contaminated.
Multiple regression relates an outcome to several predictors; logistic and ordinal logistic handle binary and ordered-categorical outcomes (a common survey shape). Mediation/moderation questions ("does X affect Y through M?") are approached through regression paths.
When your items are indicators of an underlying construct, exploratory factor analysis discovers the structure and confirmatory factor analysis tests a hypothesized one. Full structural equation modeling relates several latent variables at once; PLS-SEM suits smaller samples and prediction-focused models.
Judge a CFA/SEM by its fit indices, not the χ² alone (which almost always rejects in large samples). The common Hu & Bentler cutoffs for good fit are CFI/TLI ≥ 0.95, RMSEA ≤ 0.06, and SRMR ≤ 0.08 — report them together rather than cherry-picking one.
▸ — Repeated measures as a latent starting level plus a rate of change.
▸ — Loadings and intercepts allowed to vary with covariates — invariance, relaxed.
▸ — Composite-based SEM for smaller samples or prediction-first models.
▸ — Psychometric network: partial correlations between items, drawn as a graph.
▸ — Canonical correlation between a whole set of X and a set of Y.
▸ — Confirmatory factor analysis (test a measurement model).
▸ — Structural equation modeling (lavaan syntax).
▸ — Dimension reduction / index construction.
Reliability & agreement (Nunnally & Bernstein)
Before analyzing a scale, show it hangs together. Cronbach’s alpha and McDonald’s omega quantify internal consistency; the ICC and kappa family handle rater agreement.
▸ — Internal-consistency reliability of a scale.
▸ — Omega — reliability without alpha’s tau-equivalence assumption.
▸ — Consistency/agreement across raters or repeated measures.
▸ — Categorical agreement among 3+ raters.
Surveys and weighting (AAPOR)
Complex samples need their design weights. Weighted regression and multiple imputation for missing responses keep survey estimates honest.
▸ — Regression with survey/design weights.
▸ — Does X act on Y through M? Sobel test plus a bootstrap CI on the indirect effect.
▸ — Does the X→Y effect depend on W? Simple slopes and Johnson-Neyman.
▸ — Handle missing data by imputing + pooling.
The Bayesian alternative (Wagenmakers et al.)
Bayes factors let you quantify evidence *for* a null (not just fail to reject it) — increasingly expected in psychology. Every classical test below has a Bayesian twin.
▸ — Beta-binomial estimate for a single proportion.
▸ — Evidence for association in a contingency table.
▸ — Bayes factors for regression predictors (JZS prior).
▸ — Within-subject effects, subject treated as a random factor.
▸ — The group effect after adjusting for a covariate.
▸ — Model-averaged t-test — robust to unequal variances and heavy tails.
▸ — Bayes-factor t-test (evidence for/against a difference).
▸ — Bayes-factor ANOVA.
▸ — Bayes-factor correlation.
▸ — Test order-constrained / informative hypotheses directly.