Specialty › Complex samples

Weighted regression

Complex-samples regression fits a weighted model with Taylor-linearised standard errors that respect a stratified or clustered survey design.

What is Weighted regression?

Survey data come with sampling weights (inverse probability of selection), strata (within which sampling was independent), and clusters / PSUs (primary sampling units within which observations are correlated). Ignoring these gives biased point estimates (when ignoring weights) and incorrect SEs (when ignoring strata / clusters).

Weighted regression handles the weights; Taylor-series linearisation of the SE estimator handles the stratification / clustering. The result: design-consistent β and design-correct SEs for inference to the target population.

Kish's effective sample size = (Σw)² / Σw² quantifies the design effect: when all weights are equal, n_eff = n; with highly unequal weights, n_eff is much smaller. Report n_eff alongside n to honestly characterise precision.

When should I use Weighted regression?

  • Weighted survey data (NHANES, SHARE, ESS, etc.) where ignoring design would understate SEs and bias estimates.

What data does it need?

Response + predictors + sampling weight column + optional strata + optional cluster (PSU) + family (gaussian / binomial / poisson).

What does it report?

β + Taylor-linearised SE + Wald z / t + p + 95% CI per coefficient + Kish effective n; exponentiated β as OR (binomial) or RR (Poisson).

What does it assume?

  • Sampling weights are correct.
  • PSUs are independent across strata.

How do I interpret the result?

Compare design-based SEs to naive (unweighted) SEs to see the design effect — usually design-based are larger.

See also

References

  • Lumley (2010). Complex Surveys: A Guide to Analysis Using R.