Patterns › Clustering

Hierarchical clustering (with dendrogram)

Agglomerative hierarchical clustering builds a dendrogram from a distance matrix. Cut at a chosen height for cluster assignments.

What is Hierarchical clustering (with dendrogram)?

Start with each point as its own cluster, repeatedly merge the two closest clusters until all are merged. 'Closest' depends on the linkage method: Ward (minimise total within-SS), complete (max distance between clusters), average (mean distance), single (min distance — chains badly), etc.

Ward is the default for compact spherical clusters (similar use case to k-means). Average / UPGMA is the standard in bioinformatics for sequence trees. Single linkage's tendency to chain makes it useful for elongated clusters.

Cophenetic correlation measures how faithfully the dendrogram preserves the original distances: > 0.7 = good, < 0.5 = the linkage / distance choice is wrong for this data.

When should I use Hierarchical clustering (with dendrogram)?

  • Exploratory cluster discovery with visual merge structure.
  • When you want to pick k after seeing the dendrogram (rather than committing first).
  • Small-to-medium datasets (we cap at 2000 rows because the n² distance matrix gets unwieldy).

What data does it need?

Feature columns + linkage method + distance metric + target k for the cut.

What does it report?

SVG dendrogram coloured by cut-tree cluster + cluster sizes + per-feature means + cophenetic r.

How do I interpret the result?

Choose k by looking for big jumps in the merge heights (long vertical edges) — that indicates a natural cluster boundary.

See also

References

  • Ward (1963). Hierarchical grouping to optimize an objective function. JASA 58(301).