VerifiedEvidence: highv1.0.0

Inter-Rater Reliability

The degree of agreement between two or more independent observers assessing the same phenomenon, commonly quantified using a statistic such as Cohen's kappa.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

Concept

Theoretically, Inter-Rater Reliability is the degree of agreement between two or more independent observers assessing the same subjects using the same measurement instrument or classification system. It is founded on classical test theory and measurement reliability, ensuring that observed scores are reproducible across different raters rather than dependent on individual judgement. Inter-rater reliability exists to evaluate the consistency and objectivity of measurements involving human assessment.

Mathematically, inter-rater reliability is quantified using statistical measures of agreement appropriate to the type of data being evaluated. Cohen's kappa is commonly used for two raters and categorical data, Fleiss' kappa for multiple raters and the Intraclass Correlation Coefficient (ICC) for continuous measurements. These statistics estimate agreement beyond that expected by chance.

In practice, inter-rater reliability is assessed during the validation of clinical outcome assessments, diagnostic classifications, health technology assessments and patient-reported outcome scoring systems. Health economists evaluate inter-rater reliability to ensure consistent application of clinical classifications, resource utilisation measures and economic evaluation instruments across multiple assessors.


Purpose

Used to determine whether independent observers produce consistent measurements, supporting the reliability and reproducibility of clinical and health economic data collection.


Mathematical Formulae

Primary Formula

Cohen's kappa:

? = (P? ? P?) / (1 ? P?)

where:

P? = observed agreement

P? = expected agreement by chance

Supporting Formulae

Intraclass Correlation Coefficient (general form):

ICC = ��? / (��? + ��?)

where:

��? = between-subject variance

��? = within-subject (measurement error) variance

Percentage agreement:

Agreement (%) = (Number of agreements � Total observations) ? 100

Related Mathematical Methods

Cohen's Kappa

Fleiss' Kappa

Intraclass Correlation Coefficient

Bland-Altman Analysis

Agreement Analysis

Reliability Analysis


Example

Two clinicians independently classify 150 radiographic images as showing disease progression or no progression. They agree on 138 classifications. Cohen's kappa is calculated as ? = 0.88, indicating excellent agreement beyond chance and demonstrating high inter-rater reliability.


Excel Implementation

FunctionExample FormulaHealth Economics Application
COUNTIFS=COUNTIFS(B2:B151,C2:C151)Count identical ratings between assessors
COUNTA=COUNTA(B2:B151)Count total observations
SUMPRODUCT=SUMPRODUCT(range1,range2)Calculate agreement matrices for reliability analysis
IF=IF(D2>=0.80,"Excellent Agreement","Review Reliability")Classify inter-rater reliability
CORREL=CORREL(B2:B151,C2:C151)Assess agreement for continuous measurements

VBA (Optional)

VBA can automate agreement tables, Cohen's kappa, Intraclass Correlation Coefficient calculations and reliability reports across multiple assessors.


Sources

  • Cohen J. A Coefficient of Agreement for Nominal Scales.
  • Fleiss JL. Measuring Nominal Scale Agreement Among Many Raters.
  • Shrout PE, Fleiss JL. Intraclass Correlations: Uses in Assessing Rater Reliability.
  • Streiner DL, Norman GR, Cairney J. Health Measurement Scales: A Practical Guide to Their Development and Use.
  • COSMIN Initiative. COSMIN Methodology for Evaluating Measurement Properties.

Library

Publications

1
  • Journal article

    GRADE Guidelines: 1. Introduction — GRADE Evidence Profiles and Summary of Findings Tables — Guyatt, Oxman, Akl, Kunz, Vist, Brozek, et al., GRADE Series ed., 2011 (Journal of Clinical Epidemiology)

    The introductory paper of the GRADE (Grading of Recommendations Assessment, Development and Evaluation) series, setting out how to rate certainty of evidence (high/moderate/low/very low) and build evidence profiles and summary-of-findings tables.

Frequently Asked Questions (6)

  • What is inter-rater reliability?

    The degree of agreement between two or more independent observers assessing the same phenomenon, commonly quantified using a statistic such as Cohen's kappa.

    Source: Landis & Koch 1977

  • Why is inter-rater reliability needed when different observers assess the same thing?

    Inter-rater reliability measures how far two or more observers agree when they assess the same thing independently, and it is needed because any judgement that relies on human observation risks varying from one rater to the next. If a rating changes depending on who applied it, the score reflects the observer as much as the patient, undermining trust in the measure. Showing that different raters reach the same conclusion confirms the result stems from what is observed, not from who is looking. Agreement between raters is what makes a judgement dependable. Streiner and Norman (2008) describe this property.

    Source: Streiner & Norman 2008

  • How is inter-rater reliability measured?

    Inter-rater reliability is measured by having two or more raters independently assess the same subjects and comparing their agreement using appropriate statistics: Cohen's kappa for two raters with categorical ratings, which corrects for chance agreement; weighted kappa for ordered categories; the intraclass correlation coefficient for continuous measurements; and related measures for more raters. These express the level of agreement beyond chance. Landis and Koch proposed benchmarks for interpreting kappa values. So inter-rater reliability is quantified by statistics comparing independent raters' assessments of the same subjects, indicating how consistently the raters agree.

    Source: Landis & Koch 1977

  • Why does inter-rater reliability matter?

    Inter-rater reliability matters because when measurements or classifications involve judgement, their usefulness depends on different assessors reaching similar conclusions; low inter-rater reliability means results vary with the observer, undermining the consistency and trustworthiness of the assessment. High inter-rater reliability supports confidence that findings are reproducible and not idiosyncratic. It is important for diagnoses, ratings, and coded data used in research and practice. So inter-rater reliability matters for ensuring that assessments are consistent across raters, which is necessary for the measurements to be dependable and comparable when they rely on human judgement.

    Source: Landis & Koch 1977

  • What is Cohen's kappa in inter-rater reliability?

    Cohen's kappa is a statistic that quantifies the agreement between two raters making categorical judgements, correcting for the agreement expected by chance, and taking a value typically between zero and one, where higher values indicate stronger agreement beyond chance. Because raters may agree by chance, kappa gives a more meaningful measure than simple percentage agreement. Landis and Koch proposed benchmarks describing ranges of kappa as slight, fair, moderate, substantial, or almost perfect agreement. So Cohen's kappa is the standard measure of inter-rater agreement for categorical ratings, adjusting for chance to reflect genuine consistency between raters.

    Source: Landis & Koch 1977

  • How does inter-rater reliability differ from intra-rater reliability?

    Inter-rater reliability is the agreement between different raters assessing the same phenomenon, indicating consistency across observers, while intra-rater reliability is the agreement between repeated assessments made by the same rater on separate occasions, indicating consistency of a single rater over time. Inter-rater reliability addresses whether different people agree, and intra-rater reliability whether one person is consistent with themselves. Both use similar statistics, such as kappa or the intraclass correlation coefficient. So the two differ in whether they assess agreement between raters or the stability of a single rater's judgements.

    Source: Landis & Koch 1977

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 26 Nov 2025

Content version: 1.0.0

Canonical Identity

Term code
HE-ES-EA-025

Stable URI · Machine-readable · Resolvable · CC BY 4.0