Categorical › Logistic regression

Multiple imputation + pooled linear regression

Multiple imputation by chained equations (MICE) fills missing values repeatedly, analyses each completed dataset, and pools the results by Rubin's rules.

What is Multiple imputation + pooled linear regression?

When data are missing, listwise deletion discards information and biases the analysis if missingness depends on observed variables (MAR). MI imputes plausible values m times (each draw reflects the uncertainty about the true value), runs the analysis on each completed dataset, and pools the m estimates via Rubin's rules — combining within-imputation variance (sampling uncertainty) and between-imputation variance (missingness uncertainty).

MICE generates the m imputations by iteratively regressing each variable with missing data on the others (predictive mean matching for numeric, logistic regression for binary, proportional-odds for ordered, etc.). After enough iterations the imputations stabilise.

The fraction of missing information (FMI) per coefficient tells you how much of the SE comes from the missingness rather than the sampling. FMI > 0.2 ⇒ imputation-dominated; consider increasing m (more imputations stabilise the estimate) or rethinking the imputation model.

When should I use Multiple imputation + pooled linear regression?

  • Clinical-research workflows where listwise deletion would discard substantial data.
  • When missingness is plausibly MAR (missing at random conditional on the observed variables).
  • Pooling regression coefficients across m imputed datasets is the standard reporting.

What data does it need?

Response + predictors + # imputations m (default 5) + MICE iterations + RNG seed.

What does it report?

Per-column missingness % + the auto-selected imputation method per column; pooled β / SE / Barnard-Rubin df / p / 95% CI / FMI per coefficient.

What does it assume?

  • MAR (missing at random) — the strong assumption. MCAR is a special case where missingness is independent of everything.
  • Imputation model correctly specified.

How do I interpret the result?

m = 5 is the historical default; current recommendations are m ≈ % missing (Bodner 2008). For 30% missing, use m = 30.

If FMI > 0.5, the analysis is heavily driven by the imputation; consider whether MNAR-style sensitivity analysis is needed.

See also

References

  • Rubin (1987). Multiple Imputation for Nonresponse in Surveys.
  • Van Buuren & Groothuis-Oudshoorn (2011). Multivariate Imputation by Chained Equations. JSS 45(3).