k-means clustering partitions n observations into k clusters by minimising total within-cluster squared distance to the cluster centroid.
Iterative algorithm: pick k random centroids, assign each point to its nearest, recompute centroids as the mean of assigned points, repeat until assignments stabilise. The objective Σ_k Σ_{i ∈ C_k} ||xᵢ − μ_k||² decreases at every step.
Strengths: fast, scales to large n, gives hard cluster assignments with interpretable centroids. Weaknesses: assumes spherical equal-variance clusters (use model-based / two-step clustering for elliptical clusters), needs k pre-specified, sensitive to outliers, depends on random initialisation (run with nstart > 1 to stabilise).
Compared to hierarchical clustering: k-means is faster and works on millions of rows; hierarchical gives you a dendrogram that lets you pick k visually. Compared to DBSCAN: k-means assigns every point to a cluster (no noise class) and assumes globular shapes; DBSCAN handles arbitrary shapes and identifies outliers as noise.
≥ 2 numeric features + k + standardise toggle + nstart (random restarts).
Cluster labels per row, centroids, within-cluster sum of squares per cluster, total SS, between SS / total SS = variance explained.
Run with nstart ≥ 10 to avoid bad local minima.
Inspect per-cluster centroids on the original (un-standardised) scale to interpret what each cluster represents.