t-SNE (t-distributed stochastic neighbor embedding) reduces dimensions non-linearly while preserving local neighborhoods — the standard tool for visualising clusters.
t-SNE places each observation in 2D such that the probability of being a 'neighbour' to other points matches the same probability in high-dim space. The result is a layout where clusters in the original space stay clusters in 2D, but distances *between* clusters are not preserved — so '5cm apart' in t-SNE means nothing.
Two key knobs: perplexity controls the effective number of neighbours per point (default 30); larger perplexity ⇒ more global structure preserved at the cost of local detail. Iterations control convergence.
Compared to PCA: t-SNE handles non-linear structure that PCA misses (curved manifolds, well-separated clusters in high dim). Compared to UMAP: t-SNE has a stronger 'local-focus' bias; UMAP is faster and preserves more global structure.
Important: re-running t-SNE with a different seed gives a different 2D layout. The cluster structure is reproducible but specific point positions are not.
≥ 3 numeric columns + perplexity + optional grouping factor.
2D scatter of t-SNE coordinates, coloured by group when provided.
Do not interpret distance between clusters as meaningful — only within-cluster compactness and between-cluster separation count.
Cluster sizes in the t-SNE plot don't reflect cluster sizes in the data — t-SNE expands sparse regions and compresses dense ones.