Diagnostics › Agreement

Fleiss' kappa (3+ raters)

Fleiss' κ extends Cohen's κ to three or more raters, all classifying the same items into the same categories.

What is Fleiss' kappa (3+ raters)?

Cohen's κ pairs raters; Fleiss' κ generalises to ≥ 3 raters and reports a single overall agreement coefficient. Unlike Cohen's κ, Fleiss doesn't require the same set of raters across items — only the same number of raters per item — so it works for crowd-sourced annotation pools too.

Limitation: Fleiss treats all rater pairs symmetrically and gives one global κ; it doesn't expose pairwise disagreements. For finer analysis, run pairwise Cohen κ across all rater pairs and look for outliers.

When should I use Fleiss' kappa (3+ raters)?

  • ≥ 3 raters classifying the same items into nominal categories.
  • Crowd-sourcing reliability assessment.

What data does it need?

≥ 3 categorical columns (one per rater).

What does it report?

κ, SE, z, p, observed / expected agreement, marginal category proportions.

What does it assume?

  • Independent items.
  • Same number of raters per item.
  • Raters interchangeable (no per-rater bias is modelled).

How do I interpret the result?

Same Landis-Koch anchors as Cohen κ. Drops typically with more categories (chance agreement is harder to beat).

See also

References

  • Fleiss (1971). Measuring nominal scale agreement among many raters. Psych. Bull. 76(5).