Agglomerative hierarchical clustering builds a dendrogram from a distance matrix. Cut at a chosen height for cluster assignments.
Start with each point as its own cluster, repeatedly merge the two closest clusters until all are merged. 'Closest' depends on the linkage method: Ward (minimise total within-SS), complete (max distance between clusters), average (mean distance), single (min distance — chains badly), etc.
Ward is the default for compact spherical clusters (similar use case to k-means). Average / UPGMA is the standard in bioinformatics for sequence trees. Single linkage's tendency to chain makes it useful for elongated clusters.
Cophenetic correlation measures how faithfully the dendrogram preserves the original distances: > 0.7 = good, < 0.5 = the linkage / distance choice is wrong for this data.
Feature columns + linkage method + distance metric + target k for the cut.
SVG dendrogram coloured by cut-tree cluster + cluster sizes + per-feature means + cophenetic r.
Choose k by looking for big jumps in the merge heights (long vertical edges) — that indicates a natural cluster boundary.