Diagnostics › Agreement

Cohen's kappa (2 raters)

Cohen's κ measures agreement between exactly two raters on a categorical variable, with optional linear or quadratic weights for ordinal categories.

What is Cohen's kappa (2 raters)?

Raw percent agreement overstates agreement because some matches happen by chance — two raters classifying everyone as 'yes' agree 100% of the time but have zero discriminative agreement. Cohen's κ corrects for this: κ = (P_obs − P_exp) / (1 − P_exp), where P_exp is the agreement expected if the two raters' marginal distributions were combined independently.

κ = 1 ⇒ perfect agreement; κ = 0 ⇒ agreement no better than chance; κ < 0 ⇒ worse than chance (rare; usually a sign of inverted ratings or a coding error).

For ordinal categories (severity grades), unweighted κ treats 'mild vs severe' the same as 'mild vs moderate' — wrong. Linear weights penalise category-distance proportionally; quadratic weights penalise more harshly. Quadratic-weighted κ on ordinal data is approximately equivalent to ICC(2, 1).

When should I use Cohen's kappa (2 raters)?

  • Two raters classifying the same items into the same categories.
  • Validating a coding scheme between two annotators.
  • Switch to Fleiss κ for ≥ 3 raters, ICC for continuous ratings.

What data does it need?

Two categorical columns (one per rater) + weighting (unweighted / linear / quadratic).

What does it report?

κ, SE, 95% CI, z, p, observed / expected agreement, contingency table.

What does it assume?

  • Independent items (rows).
  • Two raters classifying the same items.
  • Marginal distributions reflect the raters' tendencies (κ adjusts for them).

Formula

κ = (P_obs − P_exp) / (1 − P_exp)

How do I interpret the result?

Landis-Koch anchors: < 0.20 slight, 0.21–0.40 fair, 0.41–0.60 moderate, 0.61–0.80 substantial, 0.81–1.00 almost perfect. Treat as calibration only; what's acceptable depends on the application.

Marginal heterogeneity (raters with very different overall rates) reduces κ even with high diagonal agreement — pair κ with raw % agreement when explaining results.

See also

References

  • Cohen (1960). A coefficient of agreement for nominal scales. EPM 20(1).
  • Landis & Koch (1977). The measurement of observer agreement for categorical data. Biometrics 33(1).