Choose the statistic from the data and design
| Feature | Cohen’s kappa | Intraclass correlation coefficient |
|---|---|---|
| Typical ratings | Nominal categories | Quantitative scores or measurements |
| Core comparison | Observed categorical agreement beyond model-expected chance agreement | Between-target variation relative to total modeled variation |
| Design information needed | Raters, categories, weighting if any, marginal distributions | Model, type, single/average measure, consistency/absolute agreement |
| Null reference often used | κ = 0 | ICC = 0, subject to model and inferential setup |
Cohen introduced kappa for agreement on nominal scales, correcting observed agreement by agreement expected under the coefficient’s chance model.1 A weighted kappa for ordered categories is a different specification and should name its weights.
An ICC is a family of coefficients, not one statistic. Koo and Li emphasize reporting the model, type, and definition—for example, whether raters are fixed or sampled, whether the target is a single rating or an average, and whether consistency or absolute agreement is required.3 Two outputs both labeled “ICC” can therefore answer different questions.
Read agreement before applying a label
For kappa, inspect the agreement table and raw percentage agreement. Kappa is affected by category prevalence and rater marginal distributions, so a seemingly high percentage agreement can coexist with a modest kappa. That is not automatically a calculation error; it reflects the coefficient’s expected-agreement correction.
For ICC, identify the full coefficient and confidence interval. A point estimate of .78 with a wide interval conveys less certainty than the same point estimate with a tight interval. If the software reports an average-measures ICC, do not interpret it as reliability of one rater.
Published kappa labels
| Kappa range | Landis and Koch label | Source/context | Category |
|---|---|---|---|
| < 0 | Poor | Observer agreement for categorical data2 | ORIGINAL CONVENTION |
| .00–.20 | Slight | Same | ORIGINAL CONVENTION |
| .21–.40 | Fair | Same | ORIGINAL CONVENTION |
| .41–.60 | Moderate | Same | ORIGINAL CONVENTION |
| .61–.80 | Substantial | Same | ORIGINAL CONVENTION |
| .81–1.00 | Almost perfect | Same | ORIGINAL CONVENTION |
These are attributed labels from Landis and Koch, not natural boundaries in the coefficient. A kappa of .60 and .61 should not lead to categorically different conclusions without substantive justification. Report the estimate, interval, category distribution, and consequences of disagreement.
Competing ICC scales
The same ICC can receive different labels under two published guidelines:
Koo and Li define < .50 poor, .50–.75 moderate, .75–.90 good, and > .90 excellent.3 Cicchetti’s psychological-assessment guideline uses < .40 poor, .40–.59 fair, .60–.74 good, and ≥ .75 excellent.4
Neither table wins by arithmetic. The report should select and cite a scale that fits its disciplinary and measurement context, then preserve the numeric estimate and interval. Where statistical thresholds come from explains why naming the source prevents a label from masquerading as a definition.
Worked comparison
Suppose an output reports ICC = .78, 95% CI [.58, .89] for single measurements under an absolute-agreement model.
- Koo and Li’s point-estimate labels call .78 “good.”
- Cicchetti’s labels call .78 “excellent.”
- The interval spans lower categories under both schemes.
- The single-measure, absolute-agreement specification is essential to the interpretation.
A defensible result therefore says more than “reliability was excellent.” It states the ICC form, point estimate, interval, selected guideline, and intended use.
Agreement is not association
Two raters can correlate strongly while differing systematically—for example, if one consistently scores every target higher. A consistency ICC may tolerate that shift; an absolute-agreement ICC is designed to count it as disagreement. This is why correlation interpretation cannot replace an agreement coefficient.
Internal consistency is another distinct job. Cronbach’s alpha concerns relationships among item scores, not whether independent raters assign the same category or measurement.
Common reading errors
- Reporting “ICC = .82” without the model, type, and definition.
- Calling a correlation coefficient inter-rater agreement.
- Applying Landis–Koch labels to ICC or ICC labels to kappa.
- Treating published bands as universal facts.
- Ignoring the confidence interval and the consequences of disagreement.
- Omitting category prevalence or the rater confusion table for kappa.
A defensible reporting sentence
“Single-measure absolute agreement was ICC = .78 (95% CI [.58, .89]). The point estimate is ‘good’ under Koo and Li’s guideline, while the interval includes values labeled moderate; the coefficient and intended use therefore matter more than the one-word category.”