Patterns › Clustering

DBSCAN

DBSCAN finds clusters of arbitrary shape from density alone, using ε (neighbourhood radius) and minPts (minimum density), and flags low-density points as noise.

What is DBSCAN?

DBSCAN groups points that are within ε of each other into 'density-reachable' clusters and labels isolated points as noise. Unlike k-means it doesn't assume spherical clusters, handles arbitrary shapes (rings, crescents), and doesn't require k.

Picking ε: plot the k-distance graph (sorted distance to the kth nearest neighbour, where k = minPts) and look for an elbow. minPts is usually 2 × dim for low-dim data, larger for high-dim.

Sensitive to scaling — always standardise unless features are already on comparable units. Fails when clusters have very different densities (one ε can't serve all).

When should I use DBSCAN?

  • Non-globular cluster shapes (rings, spirals, irregular).
  • When you want automatic noise / outlier flagging.
  • When you don't want to pre-specify k.

What data does it need?

≥ 2 numeric feature columns + ε + minPts + scaling flag.

What does it report?

Cluster label per row (−1 = noise), cluster sizes, 2D projection.

How do I interpret the result?

Most points labelled noise ⇒ ε is too small or minPts too large. One giant cluster ⇒ ε is too large.

See also

References

  • Ester, Kriegel, Sander & Xu (1996). A density-based algorithm for discovering clusters in large spatial databases with noise. KDD-96.