Regression with real-world complications: panels, endogeneity, time series, volatility.
Econometrics is regression under pressure — from serial correlation, endogeneity, unobserved heterogeneity, and non-stationarity. This guide points you at the estimator that handles each complication, plus the time-series toolkit for forecasting and dynamics.
Good practice travels with these tools. Default OLS standard errors assume homoskedastic, independent errors, so report heteroskedasticity-robust (White) or cluster-robust SEs as a matter of course. Two specification checks recur: a Hausman test to choose fixed vs random effects, and unit-root tests before regressing trending series on one another, so you don’t mistake a spurious relationship for a real one.
The regression baseline (OLS · White 1980)
Start with OLS and read its diagnostics honestly — check for the problems below before trusting the standard errors. Quantile regression describes effects away from the mean (e.g. across the wage distribution).
▸ — OLS with several predictors; the diagnostics flag the issues below.
▸ — Effects at the 10th/50th/90th percentile, not just the mean.
▸ — Lasso/ridge when predictors are many or collinear.
▸ — Outcomes censored at a floor or ceiling (top-coded income, zero spending).
▸ — Pick the family and link yourself when no preset fits.
When a regressor is correlated with the error (omitted variables, simultaneity, measurement error), OLS is biased. Instrumental-variables / two-stage least squares restores consistency with a valid instrument; propensity matching is the design-based alternative for treatment effects.
▸ — Instrumental variables / 2SLS for endogenous regressors.
▸ — Treatment-effect estimation via covariate balancing.
Panel (longitudinal) data (Hausman 1978)
With repeated observations on the same units, fixed effects sweep out time-invariant confounders while random effects are more efficient if the confounding is absent (test with Hausman logic). Dynamic panels with a lagged dependent variable need GMM.
▸ — Fixed / random / pooled panel estimators.
▸ — Arellano-Bond-style GMM for lagged-dependent-variable models.
▸ — Population-averaged effects with clustered/correlated data.
Time series: stationarity, ARIMA, forecasting (Box–Jenkins · ADF/KPSS)
Before modeling, check stationarity (ADF/KPSS) and the autocorrelation structure. ARIMA fits and forecasts univariate series; structural time-series and decomposable forecasts handle trend and seasonality more transparently.
▸ — ADF + KPSS unit-root / stationarity tests.
▸ — ACF/PACF + Ljung-Box to read the correlation structure.
▸ — Split a series into trend + seasonal + remainder before modelling it.
▸ — State-space level/slope/seasonal decomposition via a Kalman filter.
▸ — ARIMA/SARIMA fit (manual or automatic order selection).
▸ — h-step-ahead forecasts with prediction intervals.
For several interacting series, VAR captures the joint dynamics and Granger causality; if the series share a long-run equilibrium, use cointegration tests and a VECM. ARDL bounds testing handles mixed integration orders.
Financial series show volatility clustering — GARCH models the changing variance. Test whether the relationship shifts over the sample with structural-break tests.