Patterns › Dimension reduction

PCA

Principal component analysis rotates multivariate numeric data into successive orthogonal directions of maximum variance.

What is PCA?

PCA decomposes the covariance (or correlation) matrix of p numeric variables into p eigenvectors (principal components) and their eigenvalues (variance per PC). PC1 captures the largest single direction of variance; PC2 is orthogonal to PC1 and captures the next largest; and so on.

PCA is a *variance-maximising* technique — it doesn't know which variance is meaningful. If your variables are on different scales, scale them first (correlation-matrix PCA); otherwise the high-variance variables dominate the first PCs purely by unit choice.

Compared to factor analysis: PCA assumes all variance is shared (variables are deterministic linear combinations of components); EFA assumes each variable has unique variance + a small number of shared latent factors. Use PCA for dimension reduction / visualisation; use EFA for latent-construct discovery.

Compared to t-SNE / UMAP: PCA preserves global linear structure but can squash local clusters together. t-SNE / UMAP preserve local neighbourhoods at the cost of distorting global distances. Use both for visualisation when n > a few hundred.

When should I use PCA?

  • Initial multivariate visualisation — score plot of PC1 vs PC2 reveals clusters, outliers, gradients.
  • Feature reduction before regression / clustering when predictors are highly collinear.
  • Diagnostic before factor analysis: does the scree plot show a clear elbow that suggests how many factors to extract?

What data does it need?

≥ 3 numeric columns + center + scale toggles + optional grouping factor.

What does it report?

Variance explained per component (scree), Kaiser criterion (eigenvalues ≥ 1), parallel-analysis recommendation, loadings matrix, score plot of PC1 vs PC2.

What does it assume?

  • Continuous numeric variables.
  • Linearity (each PC is a linear combination of variables).
  • Scale-sensitivity: standardise first unless all variables are on the same scale.

Formula

X = U·D·V'; PC scores = X·V; variance per PC = (Dᵢ)² / (n − 1)

How do I interpret the result?

Variance explained drops as components are added; the elbow in the scree plot is the conventional cut, with Kaiser (eigenvalue ≥ 1) and parallel analysis as more principled alternatives.

Loadings are correlations between the original variable and the PC; |loading| > 0.4 is the common reporting threshold for 'this variable contributes meaningfully to this PC'.

See also

References

  • Hotelling (1933). Analysis of a complex of statistical variables into principal components. J. Educ. Psych. 24(6 & 7).