Categorical › Logistic regression

Multinomial logistic regression

Multinomial logistic regression predicts an unordered multi-category outcome, fitting one log-odds equation per non-reference category.

What is Multinomial logistic regression?

When the outcome has k > 2 unordered categories, multinomial logistic models log(P(Y = j) / P(Y = ref)) as a linear function of predictors, separately for each j ≠ ref. Each category gets its own slope vector, so the model has p · (k − 1) coefficients.

Compared to ordinal logistic, multinomial makes no ordering assumption — appropriate when categories don't have a natural rank (diagnosis types, voting choice). The cost is more parameters and lower power.

Compared to k separate binary logistic regressions: multinomial uses all the data simultaneously and gives mutually consistent probability estimates that sum to 1 across categories. One-vs-rest binaries don't.

When should I use Multinomial logistic regression?

  • Unordered categorical outcome with ≥ 3 levels (disease type, region, brand choice).
  • When you want the full set of category-vs-reference contrasts in a single model.

What data does it need?

Categorical outcome + numeric / categorical predictors. Alphabetically first level is the reference.

What does it report?

Per non-reference category: coefficient table with β, SE, z, p, OR + CI vs reference; model-level AIC + McFadden R².

What does it assume?

  • Independent observations.
  • Independence of irrelevant alternatives (IIA): adding or removing a non-reference category shouldn't change the relative odds among the others — often violated; check for plausibility.

How do I interpret the result?

An OR > 1 for category j vs reference means a one-unit increase in x multiplies the odds of being in j (vs ref) by exp(β_j). Each non-reference category has its own ORs; combining them across categories requires care.

If categories are clearly ordered, prefer ordinal logistic — fewer parameters, higher power.

See also