Principal component analysis rotates multivariate numeric data into successive orthogonal directions of maximum variance.
PCA decomposes the covariance (or correlation) matrix of p numeric variables into p eigenvectors (principal components) and their eigenvalues (variance per PC). PC1 captures the largest single direction of variance; PC2 is orthogonal to PC1 and captures the next largest; and so on.
PCA is a *variance-maximising* technique — it doesn't know which variance is meaningful. If your variables are on different scales, scale them first (correlation-matrix PCA); otherwise the high-variance variables dominate the first PCs purely by unit choice.
Compared to factor analysis: PCA assumes all variance is shared (variables are deterministic linear combinations of components); EFA assumes each variable has unique variance + a small number of shared latent factors. Use PCA for dimension reduction / visualisation; use EFA for latent-construct discovery.
Compared to t-SNE / UMAP: PCA preserves global linear structure but can squash local clusters together. t-SNE / UMAP preserve local neighbourhoods at the cost of distorting global distances. Use both for visualisation when n > a few hundred.
≥ 3 numeric columns + center + scale toggles + optional grouping factor.
Variance explained per component (scree), Kaiser criterion (eigenvalues ≥ 1), parallel-analysis recommendation, loadings matrix, score plot of PC1 vs PC2.
Variance explained drops as components are added; the elbow in the scree plot is the conventional cut, with Kaiser (eigenvalue ≥ 1) and parallel analysis as more principled alternatives.
Loadings are correlations between the original variable and the PC; |loading| > 0.4 is the common reporting threshold for 'this variable contributes meaningfully to this PC'.