Categorical › Logistic regression

Poisson regression (counts)

Generalized linear model with log link and Poisson family for non-negative integer outcomes (counts).

What is Poisson regression (counts)?

Poisson regression models log(E[Y]) as a linear function of predictors. The exponentiated β is a rate ratio — exp(β) = factor by which the expected count multiplies per unit change in the predictor. For rates with varying exposure (events per person-year), add log(exposure) as an offset via the GLM-general handler.

The Poisson distribution assumes mean = variance. Real-world counts are usually overdispersed (variance > mean) because of unmeasured heterogeneity — the residual deviance / df ratio diagnoses this. > 1.5 ⇒ try negative binomial; > 2 ⇒ definitely use negbin.

When should I use Poisson regression (counts)?

  • Count outcomes (number of events, lesions, errors).
  • Rates with fixed denominators.
  • When the count is small relative to the population at risk (else logistic / binomial is more natural).

What data does it need?

Integer-valued response + numeric predictors + intercept toggle.

What does it report?

β + SE + Wald z + p per coefficient + rate ratios = exp(β) + CI + AIC + McFadden R² + overdispersion = residual deviance / df.

What does it assume?

  • Independent observations.
  • Mean = variance (the Poisson assumption — fails for overdispersed data).
  • Log-linear relationship between predictors and log(rate).

Formula

log(E[Y_i]) = β₀ + β'X_i; E[Y_i] = exp(β₀ + β'X_i)

How do I interpret the result?

Rate ratio is the canonical effect-size. RR = 1.5 means a 50% increase in the expected count per unit of the predictor.

See also