VerifiedEvidence: highv1.0.0

Intra-Rater Reliability

The degree of agreement between repeated assessments made by the same rater evaluating the same phenomenon on separate occasions.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

Concept

Theoretically, Intra-Rater Reliability is the degree to which the same observer produces consistent measurements when assessing the same subjects on separate occasions under comparable conditions. It is founded on classical test theory and measurement reliability, evaluating the stability and repeatability of measurements made by a single assessor. Intra-rater reliability exists to determine whether observed variation reflects true changes in the measured construct rather than inconsistency in the observer.

Mathematically, intra-rater reliability is quantified using statistical measures of agreement appropriate to the measurement scale. The Intraclass Correlation Coefficient (ICC) is commonly used for continuous measurements, while Cohen's kappa is frequently applied to repeated categorical assessments. These statistics quantify the proportion of observed variability attributable to true differences between subjects rather than measurement error.

In practice, intra-rater reliability is assessed by having the same observer repeat measurements after an appropriate interval while remaining blinded to previous assessments. Health economists evaluate intra-rater reliability when validating clinical outcome assessments, health-related quality-of-life instruments, resource utilisation measures and other outcomes incorporated into economic evaluations.


Purpose

Used to determine whether repeated assessments by the same observer produce consistent measurements, supporting the reliability and reproducibility of clinical and health economic data.


Mathematical Formulae

Primary Formula

Intraclass Correlation Coefficient:

ICC = ��? / (��? + ��?)

where:

��? = between-subject variance

��? = within-subject (measurement error) variance

Supporting Formulae

Cohen's kappa:

? = (P? ? P?) / (1 ? P?)

where:

P? = observed agreement

P? = expected agreement by chance

Percentage agreement:

Agreement (%) = (Number of repeated agreements � Total observations) ? 100

Related Mathematical Methods

Intraclass Correlation Coefficient

Cohen's Kappa

Agreement Analysis

Bland-Altman Analysis

Reliability Analysis

Measurement Error Assessment


Example

A clinician independently scores the same 80 patient functional assessments on two occasions separated by four weeks while remaining blinded to the original scores. The calculated Intraclass Correlation Coefficient is 0.94, indicating excellent intra-rater reliability and demonstrating highly consistent repeated assessments.


Excel Implementation

FunctionExample FormulaHealth Economics Application
CORREL=CORREL(B2:B81,C2:C81)Assess consistency between repeated measurements
COUNTIFS=COUNTIFS(B2:B81,C2:C81)Count repeated agreements for categorical assessments
SUMPRODUCT=SUMPRODUCT(B2:B81,C2:C81)Calculate agreement matrices for repeated ratings
AVERAGE=AVERAGE(B2:B81-C2:C81)Assess average measurement differences
IF=IF(D2>=0.90,"Excellent Reliability","Further Evaluation Required")Interpret reliability results

VBA (Optional)

VBA can automate repeated measurement comparisons, Intraclass Correlation Coefficient calculations and reliability reporting for repeated observer assessments.


Sources

  • Shrout PE, Fleiss JL. Intraclass Correlations: Uses in Assessing Rater Reliability.
  • McGraw KO, Wong SP. Forming Inferences About Some Intraclass Correlation Coefficients.
  • Cohen J. A Coefficient of Agreement for Nominal Scales.
  • Streiner DL, Norman GR, Cairney J. Health Measurement Scales: A Practical Guide to Their Development and Use.
  • COSMIN Initiative. COSMIN Methodology for Evaluating Measurement Properties.

Library

Publications

1
  • BookFeatured

    Cochrane Handbook for Systematic Reviews of Interventions — Higgins, Thomas, Chandler, Cumpston, Li, Page & Welch, 2nd Edition ed., 2019 (John Wiley & Sons / Cochrane)

    The standard guide to planning, conducting, interpreting and reporting systematic reviews of health interventions, with extensive material on meta-analysis, network meta-analysis, risk of bias, GRADE, equity, complex interventions and economics evidence. Maintained as a living online resource.

Frequently Asked Questions (6)

  • What is intra-rater reliability?

    The degree of agreement between repeated assessments made by the same rater evaluating the same phenomenon on separate occasions.

    Source: Landis & Koch 1977

  • What does intra-rater reliability check about a single observer?

    Intra-rater reliability checks whether one observer, assessing the same unchanged thing on two occasions, reaches the same judgement both times. It matters because a rater whose verdict drifts from one sitting to the next introduces error even when no second observer is involved, so the same person must be consistent with themselves. Confirming this stability shows that a rating reflects the thing measured rather than the observer's shifting mood or attention. Self-consistency over time is what it establishes. Streiner and Norman (2008) describe this property.

    Source: Streiner & Norman 2008

  • How is intra-rater reliability measured?

    Intra-rater reliability is measured by having the same rater assess the same subjects on two or more separate occasions and comparing the agreement between their assessments, using statistics such as Cohen's kappa for categorical ratings, weighted kappa for ordered categories, or the intraclass correlation coefficient for continuous measurements. The occasions should be far enough apart that the rater does not simply recall the earlier judgement, but not so far that the phenomenon changes. So intra-rater reliability is quantified by comparing a single rater's repeated assessments of the same subjects, indicating how consistently that rater judges over time.

    Source: Landis & Koch 1977

  • Why does intra-rater reliability matter?

    Intra-rater reliability matters because a measurement involving judgement is only dependable if the same rater would reach the same conclusion on different occasions; low intra-rater reliability means the assessment fluctuates arbitrarily, undermining its consistency and trustworthiness. High intra-rater reliability supports confidence that a rater's measurements are stable and reproducible. It is important for assessments repeated over time, such as monitoring a patient. So intra-rater reliability matters for ensuring that a single assessor's judgements are consistent across occasions, which is necessary for measurements to be reliable when they depend on the same person's repeated evaluation.

    Source: Landis & Koch 1977

  • How is intra-rater reliability distinguished from measurement stability?

    Intra-rater reliability concerns whether the same rater gives consistent assessments of an unchanged phenomenon on separate occasions, so any disagreement reflects the rater's inconsistency rather than real change; it must be distinguished from genuine change in the phenomenon between occasions, which is not unreliability. Assessing intra-rater reliability therefore requires the phenomenon to be stable between the repeated assessments, so that differences indicate rater inconsistency. So intra-rater reliability isolates the consistency of the rater by holding the phenomenon constant, distinguishing rater-related variation from true changes in what is being measured.

    Source: Landis & Koch 1977

  • How does intra-rater reliability relate to inter-rater reliability?

    Intra-rater reliability and inter-rater reliability are complementary aspects of the reliability of judgement-based assessments: intra-rater reliability is the consistency of a single rater with themselves over repeated occasions, while inter-rater reliability is the agreement between different raters assessing the same phenomenon. Both address whether assessments are reproducible, one within a rater and the other across raters, and both use similar statistics such as kappa or the intraclass correlation coefficient. So the two together characterise the reliability of an assessment, with intra-rater reliability addressing stability over time and inter-rater reliability addressing agreement between observers.

    Source: Landis & Koch 1977

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 26 Nov 2025

Content version: 1.0.0

Canonical Identity

Term code
HE-ES-EA-026

Stable URI · Machine-readable · Resolvable · CC BY 4.0