Relate › Linear regression

Linear regression

Linear regression estimates the conditional mean of a continuous outcome as a linear function of one or more predictors.

What is Linear regression?

Linear regression models E[Y | X] = β₀ + Σ β_j · X_j. With one predictor it's the slope-intercept line through the data; with multiple it's a hyperplane in (p + 1)-space. Coefficients are estimated by ordinary least squares — minimising the sum of squared vertical distances from the line/plane to each observation.

It's the most-used model in statistics because it's interpretable (β_j is the change in Y per unit increase in X_j, holding other predictors fixed), efficient (OLS is BLUE under the Gauss-Markov conditions), and forgiving (CLT covers non-normal residuals at moderate n).

Use logistic regression for binary outcomes, Poisson / negbin for counts, Cox for time-to-event, quantile regression for non-mean conditional quantiles. Use mixed_model when observations are clustered, GEE when you want population-averaged effects with cluster-robust SEs.

When should I use Linear regression?

  • Continuous outcome modelled as a linear combination of predictors.
  • When the conditional-mean structure is approximately linear (or made so by transformations).
  • Effect-size reporting via β plus its CI; multivariable adjustment via the partial regression coefficients.

What data does it need?

Numeric response + one or more predictors (numeric or categorical via dummy coding).

What does it report?

Coefficient table with β / SE / t / p / 95% CI, R², adjusted R², overall F-test, full diagnostics (fitted vs residuals, Q-Q of residuals, Cook's distance, leverage, Breusch-Pagan + White heteroscedasticity tests, Durbin-Watson autocorrelation, VIF for multicollinearity).

What does it assume?

  • Linearity in the predictors (after any transformations).
  • Independent observations.
  • Homoscedastic residuals (constant variance vs fitted values).
  • Approximately normal residuals.
  • No high multicollinearity (VIF < 5 cautious, < 10 acceptable).

Formula

β̂ = (XᵀX)⁻¹Xᵀy, Var(β̂) = σ² · (XᵀX)⁻¹

How do I interpret the result?

Adjusted R² penalises for the number of predictors and is more honest than raw R² when comparing models with different p.

A significant Breusch-Pagan / White p means residual variance changes with the fitted value; switch to robust (sandwich) SEs or transform Y (log / sqrt). Heteroscedasticity doesn't bias β but does bias the SEs and CIs.

Durbin-Watson far from 2 ⇒ residual autocorrelation, common in time-series data; use AR(1) errors or differencing.

See also

References

  • Galton (1886). Regression towards mediocrity in hereditary stature. Journal of the Anthropological Institute 15.