The intraclass correlation coefficient measures agreement among continuous ratings from two or more raters, reporting all six Shrout-Fleiss forms with guidance on which to use.
ICC measures both consistency and absolute agreement of continuous ratings, depending on which Shrout-Fleiss form you pick. The six forms cross three model assumptions (one-way, two-way random, two-way mixed) with two units (single rater, mean of k raters).
Common choices: ICC(2, 1) for inter-rater agreement when raters are a random sample (clinical scoring by independent reviewers); ICC(3, 1) when the specific raters are fixed (always the same two pathologists); ICC(2, k) when the application is the mean of all raters (used in clinical trials with multiple readers).
ICC is also the standard reliability coefficient for test-retest with k = 2 repeated measurements per subject.
≥ 2 numeric columns (one per rater).
All six Shrout-Fleiss forms (ICC1/2/3 × single/average) with a guidance note.
Koo & Li 2016 anchors: < 0.50 poor, 0.50–0.75 moderate, 0.75–0.90 good, > 0.90 excellent reliability.
If ICC and Pearson's r differ substantially, the issue is a systematic bias (one rater scoring consistently higher than the other) that r ignores and ICC penalises — investigate the bias before reporting agreement.