UMAP embeds high-dimensional data in two dimensions non-linearly, preserving local neighbourhoods while retaining more global structure than t-SNE.
UMAP builds a weighted nearest-neighbour graph of the data in high dimensions, then optimizes a low-dimensional layout whose graph is as similar as possible. Like t-SNE it excels at revealing clusters, but it tends to preserve global relationships (distances between clusters) better and runs faster.
Two knobs matter: n_neighbours trades local detail (small) for global structure (large), and min_dist controls how tightly points pack within a cluster. As with any non-linear embedding, the axes are not interpretable and cluster sizes/inter-cluster distances should be read with caution.
A numeric matrix (rows = observations, columns = features); optionally n_neighbours and min_dist.
A 2D embedding scattered by observation, optionally coloured by a grouping variable.
Trust which points sit together far more than the exact distances between clusters or the size of a blob — these are distorted by the embedding. Re-run with different seeds/parameters to confirm a structure is real rather than an artifact, and validate any clusters with a downstream method.