Partial least squares regression reduces predictors to a small set of latent components that best predict Y. Useful when predictors outnumber observations or are highly collinear.
PLS finds the linear combinations of X that have maximum covariance with Y (vs PCA, which finds combinations with maximum variance in X alone — ignoring Y). The result is a small number of latent components that are both informative about Y and capture the structure of X.
Standard in chemometrics, spectroscopy, and genomics where p > n is the norm. The optimal number of components is picked by cross-validation: minimum RMSEP gives the best predictive model; the 1-SE rule picks the smallest model whose RMSEP is within 1 SE of minimum — more parsimonious and typically better-generalising.
Wold's VIP (variable importance in projection) gives a per-predictor importance score; predictors with VIP > 1 are 'above average' contributors and the conventional reporting threshold.
Response + ≥ 2 numeric predictors + max # components + validation method + standardise toggle.
Per-component % X-variance + % Y-variance + cumulatives; RMSEP across components; optimal-comp picks (minimum + 1-SE); β at optimum; Wold VIP per predictor.
Heavy gap between min-RMSEP and 1-SE picks ⇒ adding components doesn't help much beyond the 1-SE point; pick the simpler model.