Linear regression estimates the conditional mean of a continuous outcome as a linear function of one or more predictors.
Linear regression models E[Y | X] = β₀ + Σ β_j · X_j. With one predictor it's the slope-intercept line through the data; with multiple it's a hyperplane in (p + 1)-space. Coefficients are estimated by ordinary least squares — minimising the sum of squared vertical distances from the line/plane to each observation.
It's the most-used model in statistics because it's interpretable (β_j is the change in Y per unit increase in X_j, holding other predictors fixed), efficient (OLS is BLUE under the Gauss-Markov conditions), and forgiving (CLT covers non-normal residuals at moderate n).
Use logistic regression for binary outcomes, Poisson / negbin for counts, Cox for time-to-event, quantile regression for non-mean conditional quantiles. Use mixed_model when observations are clustered, GEE when you want population-averaged effects with cluster-robust SEs.
Numeric response + one or more predictors (numeric or categorical via dummy coding).
Coefficient table with β / SE / t / p / 95% CI, R², adjusted R², overall F-test, full diagnostics (fitted vs residuals, Q-Q of residuals, Cook's distance, leverage, Breusch-Pagan + White heteroscedasticity tests, Durbin-Watson autocorrelation, VIF for multicollinearity).
Adjusted R² penalises for the number of predictors and is more honest than raw R² when comparing models with different p.
A significant Breusch-Pagan / White p means residual variance changes with the fitted value; switch to robust (sandwich) SEs or transform Y (log / sqrt). Heteroscedasticity doesn't bias β but does bias the SEs and CIs.
Durbin-Watson far from 2 ⇒ residual autocorrelation, common in time-series data; use AR(1) errors or differencing.