Patterns › Dimension reduction

UMAP

UMAP embeds high-dimensional data in two dimensions non-linearly, preserving local neighbourhoods while retaining more global structure than t-SNE.

What is UMAP?

UMAP builds a weighted nearest-neighbour graph of the data in high dimensions, then optimizes a low-dimensional layout whose graph is as similar as possible. Like t-SNE it excels at revealing clusters, but it tends to preserve global relationships (distances between clusters) better and runs faster.

Two knobs matter: n_neighbours trades local detail (small) for global structure (large), and min_dist controls how tightly points pack within a cluster. As with any non-linear embedding, the axes are not interpretable and cluster sizes/inter-cluster distances should be read with caution.

When should I use UMAP?

  • Visualizing high-dimensional data (embeddings, omics, survey batteries) in 2D.
  • Exploratory cluster discovery where you want more global structure than t-SNE gives.
  • As a faster alternative to t-SNE on larger tables.

What data does it need?

A numeric matrix (rows = observations, columns = features); optionally n_neighbours and min_dist.

What does it report?

A 2D embedding scattered by observation, optionally coloured by a grouping variable.

What does it assume?

  • A meaningful local distance metric on the features (scale/standardize first if features differ in units).
  • Enough observations for the neighbourhood graph to be stable.

How do I interpret the result?

Trust which points sit together far more than the exact distances between clusters or the size of a blob — these are distorted by the embedding. Re-run with different seeds/parameters to confirm a structure is real rather than an artifact, and validate any clusters with a downstream method.

See also

References

  • McInnes, Healy & Melville (2018). UMAP: Uniform Manifold Approximation and Projection. arXiv:1802.03426.