Cohen’s Kappa and ICC: Interpreting Rater Agreement

Choose the statistic from the data and design

Wide data viewSwipe table horizontally
FeatureCohen’s kappaIntraclass correlation coefficient
Typical ratingsNominal categoriesQuantitative scores or measurements
Core comparisonObserved categorical agreement beyond model-expected chance agreementBetween-target variation relative to total modeled variation
Design information neededRaters, categories, weighting if any, marginal distributionsModel, type, single/average measure, consistency/absolute agreement
Null reference often usedκ = 0ICC = 0, subject to model and inferential setup

Cohen introduced kappa for agreement on nominal scales, correcting observed agreement by agreement expected under the coefficient’s chance model.1 A weighted kappa for ordered categories is a different specification and should name its weights.

An ICC is a family of coefficients, not one statistic. Koo and Li emphasize reporting the model, type, and definition—for example, whether raters are fixed or sampled, whether the target is a single rating or an average, and whether consistency or absolute agreement is required.3 Two outputs both labeled “ICC” can therefore answer different questions.

Read agreement before applying a label

For kappa, inspect the agreement table and raw percentage agreement. Kappa is affected by category prevalence and rater marginal distributions, so a seemingly high percentage agreement can coexist with a modest kappa. That is not automatically a calculation error; it reflects the coefficient’s expected-agreement correction.

For ICC, identify the full coefficient and confidence interval. A point estimate of .78 with a wide interval conveys less certainty than the same point estimate with a tight interval. If the software reports an average-measures ICC, do not interpret it as reliability of one rater.

Published kappa labels

Wide data viewSwipe table horizontally
Kappa rangeLandis and Koch labelSource/contextCategory
< 0PoorObserver agreement for categorical data2ORIGINAL CONVENTION
.00–.20SlightSameORIGINAL CONVENTION
.21–.40FairSameORIGINAL CONVENTION
.41–.60ModerateSameORIGINAL CONVENTION
.61–.80SubstantialSameORIGINAL CONVENTION
.81–1.00Almost perfectSameORIGINAL CONVENTION

These are attributed labels from Landis and Koch, not natural boundaries in the coefficient. A kappa of .60 and .61 should not lead to categorically different conclusions without substantive justification. Report the estimate, interval, category distribution, and consequences of disagreement.

Competing ICC scales

The same ICC can receive different labels under two published guidelines:

Wide data viewSwipe table horizontally
ICC rangeKoo and Li (2016)Cicchetti (1994)Provenance
< .40PoorPoorDifferent guideline contexts3 4
.40–.49PoorFairCompeting labels
.50–.59ModerateFairCompeting labels
.60–.74ModerateGoodCompeting labels
.75–.90Good, with >.90 excellentExcellentCompeting labels

Koo and Li define < .50 poor, .50–.75 moderate, .75–.90 good, and > .90 excellent.3 Cicchetti’s psychological-assessment guideline uses < .40 poor, .40–.59 fair, .60–.74 good, and ≥ .75 excellent.4

Neither table wins by arithmetic. The report should select and cite a scale that fits its disciplinary and measurement context, then preserve the numeric estimate and interval. Where statistical thresholds come from explains why naming the source prevents a label from masquerading as a definition.

Worked comparison

Suppose an output reports ICC = .78, 95% CI [.58, .89] for single measurements under an absolute-agreement model.

  • Koo and Li’s point-estimate labels call .78 “good.”
  • Cicchetti’s labels call .78 “excellent.”
  • The interval spans lower categories under both schemes.
  • The single-measure, absolute-agreement specification is essential to the interpretation.

A defensible result therefore says more than “reliability was excellent.” It states the ICC form, point estimate, interval, selected guideline, and intended use.

Agreement is not association

Two raters can correlate strongly while differing systematically—for example, if one consistently scores every target higher. A consistency ICC may tolerate that shift; an absolute-agreement ICC is designed to count it as disagreement. This is why correlation interpretation cannot replace an agreement coefficient.

Internal consistency is another distinct job. Cronbach’s alpha concerns relationships among item scores, not whether independent raters assign the same category or measurement.

Common reading errors

  • Reporting “ICC = .82” without the model, type, and definition.
  • Calling a correlation coefficient inter-rater agreement.
  • Applying Landis–Koch labels to ICC or ICC labels to kappa.
  • Treating published bands as universal facts.
  • Ignoring the confidence interval and the consequences of disagreement.
  • Omitting category prevalence or the rater confusion table for kappa.

A defensible reporting sentence

“Single-measure absolute agreement was ICC = .78 (95% CI [.58, .89]). The point estimate is ‘good’ under Koo and Li’s guideline, while the interval includes values labeled moderate; the coefficient and intended use therefore matter more than the one-word category.”

Evidence trail

Sources and notes

  1. Definition and worked agreement tables
  2. J. Richard Landis and Gary G. Koch (1977). The Measurement of Observer Agreement for Categorical Data
    Table of strength-of-agreement labels
  3. Interpretation guideline and model/type/definition sections
  4. Reliability guideline table

Optional analytics help us understand site use. They remain off unless you accept.

Your preferences