Binary logistic regression models log-odds of a 0/1 outcome as a linear function of predictors. Effects are reported as odds ratios.
Logistic regression is the workhorse for binary outcomes. The link function logit(p) = log(p/(1−p)) maps probabilities to the real line, so log-odds can be modelled linearly without constraining predicted probabilities to [0, 1]. Each coefficient β_j is the log of the odds ratio for a one-unit change in x_j; exp(β_j) is the odds ratio itself.
Compared to linear regression of a 0/1 outcome: linear models can produce predictions outside [0, 1] and have heteroscedastic residuals (variance depends on p). Logistic regression sidesteps both issues. For probabilities near 0.5, the two often agree numerically; near 0 or 1 they diverge.
Three close relatives: ordinal_logistic for ordered multi-category outcomes (proportional-odds model), multinomial_logistic for unordered multi-category outcomes (one log-odds equation per non-reference category), and cond_logistic for matched case-control studies (eliminates the within-stratum nuisance intercept).
Binary outcome + numeric / categorical predictors + intercept toggle.
β (log-odds), SE, z, p, OR + 95% CI per predictor; AIC; McFadden pseudo-R²; predicted-probability score for ROC.
OR > 1 ⇒ predictor increases the odds of the outcome; OR < 1 ⇒ decreases. The point estimate matters far more than statistical significance; an OR of 1.05 with p = 0.001 in a huge sample is rarely clinically meaningful.
OR ≠ risk ratio: with a common outcome (incidence > 10%), the OR overstates the RR. Switch to Poisson with log link or use a log-binomial model when the outcome is common and you want RRs directly.
Pseudo-R²s are not directly comparable to OLS R²s — McFadden values of 0.2 − 0.4 indicate excellent fit, not poor.