Concept Architecture
Concept
Theoretically, Test-Retest Reliability is the degree to which a measurement instrument produces consistent results when administered repeatedly to the same individuals under equivalent conditions over a period during which the underlying construct is expected to remain unchanged. It is founded on classical test theory and measurement reliability, evaluating the temporal stability of an instrument. Test-retest reliability exists to determine whether repeated measurements reflect true stability rather than measurement error or random variation.
Mathematically, test-retest reliability is quantified by measuring the agreement between repeated observations obtained from the same subjects. The Intraclass Correlation Coefficient (ICC) is the most commonly recommended statistic for continuous measurements, while Cohen's kappa is frequently used for categorical outcomes. Correlation coefficients alone assess association but do not necessarily measure agreement.
In practice, test-retest reliability is assessed by administering the same instrument to the same participants after an appropriate interval during which no meaningful clinical change is expected. Health economists evaluate test-retest reliability when validating health-related quality-of-life instruments, utility measures and patient-reported outcome measures used in health technology assessment and economic evaluation.
Purpose
Used to evaluate the temporal stability of measurement instruments and ensure that repeated assessments produce consistent results when the underlying construct has not changed.
Mathematical Formulae
Primary Formula
Intraclass Correlation Coefficient:
ICC = ��? / (��? + ��?)
where:
��? = between-subject variance
��? = within-subject (measurement error) variance
Supporting Formulae
Cohen's kappa:
? = (P? ? P?) / (1 ? P?)
Pearson correlation coefficient:
r = Cov(X,Y) / (�?�?)
Related Mathematical Methods
Intraclass Correlation Coefficient
Cohen's Kappa
Bland-Altman Analysis
Measurement Error Assessment
Internal Consistency Reliability
Inter-Rater Reliability
Intra-Rater Reliability
Example
A health-related quality-of-life questionnaire is administered to 120 patients with stable chronic obstructive pulmonary disease on two occasions separated by two weeks. The Intraclass Correlation Coefficient is 0.93, indicating excellent agreement between repeated assessments and demonstrating that the questionnaire provides highly stable measurements over time.
Excel Implementation
| Function | Example Formula | Health Economics Application |
|---|---|---|
| CORREL | =CORREL(B2:B121,C2:C121) | Assess association between repeated measurements |
| AVERAGE | =AVERAGE(B2:B121-C2:C121) | Calculate the mean difference between test and retest scores |
| STDEV.S | =STDEV.S(B2:B121-C2:C121) | Assess variability of repeated measurement differences |
| ABS | =ABS(B2-C2) | Calculate absolute differences between repeated observations |
| IF | =IF(D2>=0.90,"Excellent Reliability","Further Evaluation Required") | Interpret test-retest reliability statistics |
VBA (Optional)
VBA can automate repeated measurement analyses, calculate agreement statistics and generate reliability reports for health outcome instruments.
Sources
- Streiner DL, Norman GR, Cairney J. Health Measurement Scales: A Practical Guide to Their Development and Use.
- Shrout PE, Fleiss JL. Intraclass Correlations: Uses in Assessing Rater Reliability.
- McGraw KO, Wong SP. Forming Inferences About Some Intraclass Correlation Coefficients.
- COSMIN Initiative. COSMIN Methodology for Evaluating Measurement Properties.
- DeVellis RF. Scale Development: Theory and Applications.
Related Concepts (2)
Library
Publications
1
GRADE Guidelines: 1. Introduction — GRADE Evidence Profiles and Summary of Findings Tables — Guyatt, Oxman, Akl, Kunz, Vist, Brozek, et al., GRADE Series ed., 2011 (Journal of Clinical Epidemiology)
The introductory paper of the GRADE (Grading of Recommendations Assessment, Development and Evaluation) series, setting out how to rate certainty of evidence (high/moderate/low/very low) and build evidence profiles and summary-of-findings tables.
Journal ArticleView source →
Frequently Asked Questions (6)
What is test-retest reliability?
The degree of agreement between scores from the same instrument given to the same people on two separate occasions, when no real change is expected.
Source: Landis & Koch 1977
What does test-retest reliability tell us about a stable trait?
Test-retest reliability tells us whether an instrument gives the same score when the same people are measured again after a gap during which nothing about them has really changed. For a stable trait, close agreement between the two occasions shows the score reflects the person rather than the moment, so day-to-day fluctuation in the instrument is small. The interval must be long enough to prevent simple recall yet short enough that no true change intervenes. Stability over time is what it demonstrates. Streiner and Norman (2008) describe this property.
Source: Streiner & Norman 2008
How is test-retest reliability measured?
Test-retest reliability is measured by administering the same instrument to the same people on two occasions separated by an interval during which no real change is expected, and comparing the scores, typically using the intraclass correlation coefficient for continuous measures or kappa for categorical ones. Higher agreement indicates greater test-retest reliability. The interval must be chosen carefully: long enough that respondents do not simply recall their previous answers, but short enough that the condition does not genuinely change. So test-retest reliability is quantified by the correlation between repeated administrations under conditions of no expected change.
Source: Landis & Koch 1977
Why is test-retest reliability important?
Test-retest reliability is important because an instrument used to measure a stable characteristic, or to detect change over time, must give consistent results when the underlying condition is unchanged; if scores vary arbitrarily between occasions, the instrument is unreliable, and apparent changes could reflect measurement instability rather than real change. High test-retest reliability supports confidence that observed differences over time are genuine. So test-retest reliability matters for ensuring that an instrument's measurements are stable over time, which is necessary for interpreting scores and for distinguishing real change from measurement error in longitudinal use.
Source: Landis & Koch 1977
What interval should be used for test-retest reliability?
The interval for test-retest reliability should be chosen so that no real change in the underlying condition is expected between administrations, yet long enough that respondents do not simply recall and repeat their previous answers, which would inflate the apparent reliability. A very short interval risks recall effects, while a long one risks genuine change occurring. The appropriate interval depends on the stability of the characteristic being measured. So selecting the interval involves balancing the avoidance of recall against the avoidance of true change, so that agreement between occasions reflects the instrument's stability rather than memory or real change.
Source: Landis & Koch 1977
How does test-retest reliability differ from internal consistency?
Test-retest reliability differs from internal consistency in what they assess: test-retest reliability concerns the stability of an instrument's scores across separate occasions, examining consistency over time, while internal consistency concerns the coherence of the instrument's items within a single administration, examining whether the items measure the same construct. Test-retest reliability is measured by correlating repeated administrations, and internal consistency by statistics such as Cronbach's alpha from item intercorrelations. So the two address different aspects of reliability: consistency over time versus consistency among items, both contributing to the overall reliability of an instrument.
Source: Cronbach 1951
Trust Record
Verified by Dr Darrin Baines
British health economist
Professional identity: darrinbaines.org
Verification date: 26 Nov 2025
Content version: 1.0.0
Canonical Identity
- Term code
- HE-ES-EA-040
Stable URI · Machine-readable · Resolvable · CC BY 4.0