Categorical › Logistic regression

GEE (generalized estimating equations)

Generalized estimating equations fit population-averaged regressions for clustered or longitudinal binary, count or continuous outcomes.

What is GEE (generalized estimating equations)?

GEE estimates *population-averaged* effects (also called marginal effects): how much does the mean of Y change in the population when X changes by 1? This contrasts with mixed-effects models, which estimate *subject-specific* effects — change in Y *for the same subject* when X changes. For Gaussian outcomes the two coincide; for binary or count outcomes they don't, and the choice depends on your scientific question.

GEE specifies a working correlation structure (independence / exchangeable / AR(1) / unstructured) but uses Liang-Zeger sandwich SEs that are consistent even when the working correlation is wrong. So you don't need to get the correlation structure right for valid inference — though a closer structure gives more efficient estimates. QIC chooses among structures.

Population-averaged framing is the right choice when the question is at the population level (public health policy effects, marketing reach) or when subject-specific effects are unidentifiable. Use mixed models when you care about between- vs within-subject variance components or about predicting trajectories for specific subjects.

When should I use GEE (generalized estimating equations)?

  • Clustered or longitudinal data with non-Gaussian outcomes where you want population-averaged effects.
  • When the within-cluster correlation isn't of scientific interest but you need to account for it in SEs.
  • Robustness-first analyses: sandwich SEs make GEE less sensitive to model misspecification than mixed models.

What data does it need?

Response + predictors + subject / cluster ID + family (gaussian / binomial / poisson / Gamma) + working correlation.

What does it report?

β + sandwich SE + Wald χ² + p + 95% CI per coefficient + working-correlation α + QIC for structure selection.

What does it assume?

  • MCAR or MAR missingness (GEE is not consistent under MNAR).
  • Correctly specified mean structure.
  • Independence between clusters.

How do I interpret the result?

β estimates are population-averaged; for binary outcomes they are smaller in absolute value than subject-specific (mixed-model) estimates by a factor depending on the random-intercept variance. Don't directly compare GEE β to mixed-model β.

See also

References

  • Liang & Zeger (1986). Longitudinal data analysis using generalized linear models. Biometrika 73(1).