Fleiss' κ extends Cohen's κ to three or more raters, all classifying the same items into the same categories.
Cohen's κ pairs raters; Fleiss' κ generalises to ≥ 3 raters and reports a single overall agreement coefficient. Unlike Cohen's κ, Fleiss doesn't require the same set of raters across items — only the same number of raters per item — so it works for crowd-sourced annotation pools too.
Limitation: Fleiss treats all rater pairs symmetrically and gives one global κ; it doesn't expose pairwise disagreements. For finer analysis, run pairwise Cohen κ across all rater pairs and look for outliers.
≥ 3 categorical columns (one per rater).
κ, SE, z, p, observed / expected agreement, marginal category proportions.
Same Landis-Koch anchors as Cohen κ. Drops typically with more categories (chance agreement is harder to beat).