Definition
A statistical index that quantifies the proportion of total variance in continuous or ordinal ratings attributable to differences between target units (e.g., subjects, items, groups) rather than to measurement error or rater variability, conditional on a specified model of effects.

Principle

Principle
The ICC decomposes observed variance into components (between‑unit and within‑unit/error); values closer to 1 indicate that most variance is between targets and thus measurements are more reproducible across raters or occasions under the chosen model.

Demonstration

Demonstration
Illustrative scenario: two interviewers independently rate religious commitment on a continuous scale for 50 participants. Recognition: compute ICC using a two‑way model. Action: a high ICC supports treating participant scores as stable characteristics; a low ICC prompts examiner training or instrument revision. Consequence: the ICC guides whether to model responses as clustered and whether to average raters' scores.

Misapplication

Misapplication
Applying an ICC formula without matching the model (one‑way vs two‑way; consistency vs absolute agreement) to the study design, or using ICC on nominal data—both produce misleading reliability estimates.

Consequence

Consequence
Correct selection and interpretation of ICC informs measurement decisions (e.g., averaging raters, estimating effective sample size, multilevel modeling). Misuse leads to incorrect conclusions about reliability, biased standard errors, and improper analytical choices.

Reversal

Reversal
When assumptions underlying variance decomposition are violated (heteroscedasticity, nonlinearity, nonindependence) or when data are strongly nonnormal or categorical, ICC estimates can be biased and alternative indices or robust methods should be used.

Boundary

Boundary
Clearly within: continuous or approximately continuous ratings from multiple raters or occasions with a specified random/ fixed effects structure. Boundary case: ordinal Likert scales—ICC can be used with care or with polychoric methods. Clearly outside: nominal category assignments—use kappa statistics instead.

Semantic Tension

Semantic Tension
ICC ↔ Kappa/Weighted Kappa — ICC addresses variance decomposition for continuous/ordinal measures and clustering; kappa is appropriate for nominal agreement and accounts differently for chance agreement.

Synthesis

Synthesis
The ICC links reliability to the proportionate contribution of between‑unit variance in a chosen statistical model; its utility depends on model choice and data scale, so it functions as both a reliability metric and a diagnostic for analytic strategy.