Cohen's κ measures agreement between exactly two raters on a categorical variable, with optional linear or quadratic weights for ordinal categories.
Raw percent agreement overstates agreement because some matches happen by chance — two raters classifying everyone as 'yes' agree 100% of the time but have zero discriminative agreement. Cohen's κ corrects for this: κ = (P_obs − P_exp) / (1 − P_exp), where P_exp is the agreement expected if the two raters' marginal distributions were combined independently.
κ = 1 ⇒ perfect agreement; κ = 0 ⇒ agreement no better than chance; κ < 0 ⇒ worse than chance (rare; usually a sign of inverted ratings or a coding error).
For ordinal categories (severity grades), unweighted κ treats 'mild vs severe' the same as 'mild vs moderate' — wrong. Linear weights penalise category-distance proportionally; quadratic weights penalise more harshly. Quadratic-weighted κ on ordinal data is approximately equivalent to ICC(2, 1).
Two categorical columns (one per rater) + weighting (unweighted / linear / quadratic).
κ, SE, 95% CI, z, p, observed / expected agreement, contingency table.
Landis-Koch anchors: < 0.20 slight, 0.21–0.40 fair, 0.41–0.60 moderate, 0.61–0.80 substantial, 0.81–1.00 almost perfect. Treat as calibration only; what's acceptable depends on the application.
Marginal heterogeneity (raters with very different overall rates) reduces κ even with high diagonal agreement — pair κ with raw % agreement when explaining results.