Patterns › Clustering

TwoStep cluster (BIC-selected k)

TwoStep-style cluster analysis evaluates 14 Gaussian-mixture covariance structures across k = 1..maxK and picks the (k, family) pair with the highest BIC.

What is TwoStep cluster (BIC-selected k)?

Model-based clustering treats each cluster as a Gaussian distribution. The model evaluates 14 different covariance assumptions (spherical / diagonal / ellipsoidal × equal-volume / variable-volume / equal-orientation / etc.) at each k, fits each via EM, and picks the (k, family) maximising BIC.

Advantage over k-means / hierarchical: principled k selection by an information criterion rather than visual elbow / dendrogram inspection. Disadvantage: assumes clusters are Gaussian — fails for clearly non-Gaussian shapes (use DBSCAN there).

Other two-step implementations use a slightly different distance metric but the same BIC-based selection — the headline output (auto-k + cluster profiles) is comparable for the continuous-feature case.

When should I use TwoStep cluster (BIC-selected k)?

  • When the right k isn't obvious and you want a defensible auto-selection.
  • Reasonable cluster shapes (Gaussian-ish).
  • Continuous features only in v1; categorical extensions are a separate workflow.

What data does it need?

Feature columns + max k + standardise toggle.

What does it report?

Selected k + family + BIC + cluster sizes + per-feature means + BIC curve across k + PCA scatter coloured by cluster.

How do I interpret the result?

Sharp BIC peak ⇒ confident k. Plateau ⇒ k is loosely determined; try alternative methods.

See also

References

  • Fraley & Raftery (2002). Model-based clustering, discriminant analysis, and density estimation. JASA 97(458).