Patterns › Clustering

Fuzzy c-means clustering

Fuzzy c-means clustering gives every observation a membership degree in every cluster, replacing the hard labels of k-means with soft assignments.

What is Fuzzy c-means clustering?

FCM minimizes a membership-weighted within-cluster distance, with the fuzzifier m controlling softness: m → 1 recovers crisp k-means behaviour, larger m spreads membership across clusters. Hard labels (for the plot and sizes) take each observation's maximum membership.

Dunn's partition coefficient (Σu²/n) summarizes decisiveness: 1/k means totally fuzzy (uniform memberships), 1 means fully crisp. Mean max-membership reads similarly on a per-observation basis.

Useful when cluster boundaries are genuinely gradual — observations near a boundary get split membership instead of an arbitrary hard label.

When should I use Fuzzy c-means clustering?

  • Overlapping / gradual cluster structure where hard assignment misrepresents boundary cases.
  • As a robustness check on k-means: if the partition coefficient is near 1/k, the 'clusters' may not be real.

What data does it need?

Numeric feature columns + number of clusters k + fuzzifier m (2 = standard) + standardize toggle.

What does it report?

Cluster sizes (by max membership), partition coefficient, mean max-membership, PCA 2D scatter colored by cluster.

What does it assume?

  • Roughly spherical clusters in (standardized) feature space, like k-means.
  • k chosen by the user — try the k-means elbow first.

How do I interpret the result?

Partition coefficient near 1: crisp, k-means would say the same. Near 1/k: memberships are near-uniform — reconsider k or whether clusters exist.

See also

References

  • Bezdek (1981). Pattern Recognition with Fuzzy Objective Function Algorithms.