VerifiedEvidence: highv1.0.0

Reliability Assessment

An evaluation of the consistency and reproducibility of an instrument's results, covering internal consistency, inter-rater, and test-retest reliability.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

How reliability assessment evaluates score consistency

Reliability assessment examines whether an instrument produces sufficiently consistent and reproducible scores for a defined population, setting, and use. It separates variation attributable to real differences among people from variation created by items, raters, occasions, devices, or other measurement conditions. This page explains the main reliability forms, the statistics used to estimate them, and how reliability affects interpretation, sample size, and economic modelling.

Reliability is conditional rather than permanent

An instrument is not simply reliable or unreliable in every context. Reliability depends on the population's variability, the raters, administration mode, timing, language, setting, scoring rule, and intended decision. Evidence from one use should not be transferred automatically to another.

A reliability claim should identify:

  • The instrument, version, domain, and score.
  • The target population and range of severity.
  • The administration mode, setting, and language.
  • The raters, devices, or occasions being compared.
  • The interval between measurements.
  • Whether scores inform group comparisons or individual decisions.
  • The form and threshold of reliability required.

Reliability and validity answer different questions

Reliability asks whether measurement is consistent, while validity asks whether the score supports the intended interpretation and use. A measure can consistently produce the wrong value and therefore be reliable but invalid. Low reliability limits validity because unstable scores cannot support precise inference.

Measurement propertyMain questionExample
ReliabilityAre scores sufficiently consistent across relevant replications?Do stable patients receive similar scores one week apart?
AgreementHow close are repeated scores in the original units?Do two raters differ by no more than two points?
ValidityDoes the score measure and support the intended interpretation?Does the scale represent fatigue in the target population?
ResponsivenessCan the score detect change when the concept changes?Does improvement after effective treatment produce score change?
InterpretabilityWhat does a score or change mean?Is a five-point improvement meaningful to patients?

Classical test theory separates signal and error

Classical test theory represents an observed score as a true score plus measurement error. The true score is the expected score across hypothetical replications under the defined conditions, not an independently observable perfect value. Reliability is the proportion of observed variance attributable to true-score variance.

$$ X=T+E $$

$$ Reliability=\frac{Var(T)}{Var(X)}=1-\frac{Var(E)}{Var(X)} $$

Reliability can be higher in a heterogeneous sample because between-person variance is larger, even when absolute measurement error is unchanged. Sample composition should therefore accompany every coefficient.

Internal consistency evaluates related items

Internal consistency assesses whether items intended to measure the same construct produce coherent responses. It is relevant to multi-item reflective scales, not to checklists whose components need not correlate. High internal consistency does not prove unidimensionality, content validity, or absence of redundant items.

Before estimating a coefficient, reviewers should confirm that:

  • Items are intended as indicators of one underlying construct or domain.
  • The scoring rule is supported by structural-validity evidence.
  • Items are not duplicated or locally dependent.
  • The response distribution is appropriate for the method.
  • The sample represents the intended population.

Cronbach's alpha has restrictive assumptions

Cronbach's alpha is widely reported as an internal-consistency coefficient. It is influenced by the number of items and average inter-item covariance. Its usual reliability interpretation is strongest when items are essentially tau-equivalent and errors satisfy relevant assumptions.

$$ \alpha=\frac{k}{k-1}\left(1-\frac{\sum_{i=1}^{k}\sigma_i^2}{\sigma_X^2}\right) $$

where (k) is the number of items, (\sigma_i^2) is item variance, and (\sigma_X^2) is total-score variance. Adding similar items can raise alpha without improving construct coverage.

Omega can better accommodate unequal item loadings

McDonald's omega uses a factor model and can provide a more appropriate reliability estimate when item loadings differ. Its credibility depends on the fitted dimensional structure and error assumptions. Reporting omega does not remove the need to examine model fit, local dependence, and cross-loadings.

For a simple congeneric one-factor model, coefficient omega total can be represented as:

$$ \omega=\frac{\left(\sum_i\lambda_i\right)^2}{\left(\sum_i\lambda_i\right)^2+\sum_i\theta_i} $$

where (\lambda_i) are item loadings and (\theta_i) are error variances under the specified model. The exact form should match the scoring and factor structure.

Very high internal consistency can be a warning

A very high coefficient may indicate that items repeat nearly the same content, increasing respondent burden without adding useful information. It can also reflect a narrow construct or method effects. Item-total relationships, content coverage, factor structure, and information across the trait range should be reviewed.

No universal alpha or omega threshold determines acceptability. Individual decision making generally requires greater precision than group-level research, but context and consequences matter.

Test-retest reliability assesses stability over time

Test-retest reliability evaluates whether scores remain consistent when the underlying concept is expected to be stable. The interval should be long enough to reduce memory and short enough to avoid real change. Treatment, learning, fatigue, disease fluctuation, and life events can all violate stability.

The study should define how stability was established rather than assume it from elapsed time. Patient or clinician global assessments, unchanged treatment, and clinical criteria can support the stable-sample definition, but each has limitations.

Inter-rater reliability assesses different assessors

Inter-rater reliability evaluates whether different raters produce consistent scores for the same subjects or material. It is relevant to clinician-reported outcomes, diagnostic ratings, imaging, chart abstraction, and coded observations. Raters should represent those who will use the instrument in practice.

Design considerations include:

  • Rater training, experience, and certification.
  • Independent and blinded assessments.
  • The same information and observation period for each rater.
  • Balanced assignment of raters and subjects.
  • Sufficient cases across the severity range.
  • Whether raters are fixed or sampled from a wider population.

Intra-rater reliability assesses one rater over repetitions

Intra-rater reliability evaluates consistency when the same rater repeats an assessment under comparable conditions. The interval must reduce recall while keeping the underlying material stable. Repeated review of static images differs from repeated clinical examination because the patient can change.

Learning or drift can change ratings over time. Ongoing calibration is important when an instrument is used in long studies or routine programmes.

Intraclass correlation coefficients have several forms

The intraclass correlation coefficient, or ICC, estimates the proportion of variance attributable to differences among subjects relative to total variance. Different ICC models answer different questions about raters, agreement, consistency, and single versus average ratings. Reporting only “ICC” is incomplete.

A variance-component form is:

$$ ICC=\frac{\sigma^2_{subject}}{\sigma^2_{subject}+\sigma^2_{error}} $$

More complex designs can include rater, occasion, and interaction variance. The selected form should state whether raters are fixed or random and whether absolute agreement or consistency is required.

Consistency and absolute agreement are not the same

Consistency reliability can be high when raters preserve rank order despite one systematically scoring higher. Absolute-agreement reliability treats systematic rater differences as disagreement. Clinical interchangeability usually requires absolute agreement, while some research questions may focus on ranking.

The choice should follow intended use and be declared before results. Selecting the larger coefficient after analysis overstates performance.

Kappa evaluates agreement for categorical ratings

Cohen's kappa measures agreement between two raters beyond the amount expected by chance under its model. Weighted kappa can assign partial agreement to ordered categories. Multi-rater variants are needed when more than two raters contribute.

$$ \kappa=\frac{P_o-P_e}{1-P_e} $$

where (P_o) is observed agreement and (P_e) is chance-expected agreement. Kappa can be low despite high observed agreement when prevalence is extreme or marginal distributions differ, so the confusion matrix and raw agreement should also be reported.

Weighted kappa requires defensible weights

For ordered categories, disagreements between adjacent categories may be less serious than disagreements across the full scale. Linear or quadratic weights encode that judgment. Different weights can materially change the result and should be specified and justified.

Weighted kappa can resemble an ICC under some conditions, but the measures are not interchangeable without examining scale, design, and assumptions.

Reliability and agreement should be reported together

Reliability is a relative property that compares measurement error with differences among subjects. Agreement quantifies error in the original units. Two samples can have different reliability despite similar absolute error because their between-person variability differs.

Agreement can be described through:

  • Mean difference between repeated measures.
  • Standard deviation of the differences.
  • Limits of agreement.
  • Standard error of measurement.
  • Smallest detectable change.
  • Proportion of exact or clinically acceptable agreement.

Bland-Altman analysis shows repeated-measure differences

A Bland-Altman plot displays the difference between two measurements against their mean. It can reveal systematic bias, changing variability, outliers, and whether disagreement depends on the magnitude of measurement. It is more informative than correlation for assessing interchangeability.

If differences (d_i=X_{i2}-X_{i1}) are approximately normally distributed, 95% limits of agreement are:

$$ \overline{d}\pm1.96\times SD(d) $$

The limits should be compared with a clinically acceptable disagreement range. Repeated observations and heteroscedastic differences require adapted methods.

Correlation is not reliability or agreement

Pearson or Spearman correlation measures association, not equality of repeated values. Two raters can have perfect correlation when one always scores ten points higher. Correlation also ignores systematic bias and can be inflated by a wide score range.

Scatterplots and correlation can supplement analysis, but they should not replace ICC, kappa, or agreement methods selected for the measurement scale and design.

The standard error of measurement quantifies score error

The standard error of measurement, or SEM, expresses measurement error in the instrument's units. Under classical assumptions, it can be estimated from score variability and reliability. It supports interpretation of individual scores and change.

$$ SEM=SD\sqrt{1-r} $$

where (r) is the relevant reliability coefficient. The SD and reliability must come from an appropriate population and design; mixing values from different studies can produce an incoherent estimate.

Detectable change differs from meaningful change

The smallest detectable change estimates how much change is needed to exceed expected measurement error at a specified confidence level. It does not establish that patients consider the change important. Meaningful change requires anchor-based and content-validity evidence.

For two measurements and approximately 95% confidence:

$$ SDC_{95}=1.96\sqrt{2}\times SEM $$

If a meaningful-change threshold is smaller than the detectable change for an individual, the instrument may be unable to distinguish that meaningful change reliably at the individual level.

Reliability affects observed treatment effects

Measurement error can attenuate associations and increase variance. In randomised trials, non-differential outcome error can reduce power even when it does not create systematic treatment-group bias. Differential measurement caused by unblinded assessment can introduce bias as well as unreliability.

Improving reliability can reduce required sample size or increase precision, but a highly reliable measure of the wrong concept remains inappropriate. Measurement selection should balance reliability, validity, responsiveness, interpretability, and burden.

Reliability affects sample-size planning

Lower reliability reduces the proportion of observed variance that represents signal. Under simplified assumptions, correcting an observed correlation for attenuation uses the reliabilities of both measures:

$$ \rho_{XY,true}=\frac{\rho_{XY,observed}}{\sqrt{r_Xr_Y}} $$

This formula depends on classical independent errors and should not be used mechanically. Planning should use simulation or design-specific formulas when outcomes, clustering, missingness, and repeated measures are complex.

Item-response theory provides conditional precision

Item-response theory and Rasch models estimate information across levels of the underlying trait. Reliability can vary across the score range even when a single marginal coefficient looks acceptable. Conditional standard errors show where the instrument is most and least precise.

Computerised adaptive testing can select informative items for each respondent, reducing burden while maintaining precision. Its item bank, calibration, stopping rules, and comparability require validation in the intended population.

Generalisability theory separates several error sources

Generalisability theory extends classical reliability by estimating variance from facets such as raters, items, occasions, and their interactions. A generalisability study identifies important error sources, while a decision study predicts reliability under alternative numbers of items or raters. This can guide efficient assessment design.

A broad decomposition is:

$$ \sigma^2_{observed}=\sigma^2_{person}+\sigma^2_{item}+\sigma^2_{rater}+\sigma^2_{occasion}+\sigma^2_{interactions}+\sigma^2_{residual} $$

The relevant coefficient depends on whether the intended interpretation is relative ranking or absolute decision making.

Rater training can improve but not guarantee reliability

Manuals, examples, certification, and calibration can reduce rater variation. High agreement during training may decline in routine use because of drift, staff turnover, workload, or case complexity. Monitoring and retraining should be planned for long-running studies or services.

Training can also make raters consistently apply a flawed rule. Validity and clinical relevance remain separate requirements.

Automation changes the source of measurement error

Devices and algorithms can remove some human variation while introducing calibration, hardware, software, signal-processing, and data-quality errors. Repeated measurements from the same device do not establish agreement across devices, versions, sites, or populations. Algorithm updates can change the score even when the instrument name remains constant.

Reliability assessment should record device model, software version, settings, operator, environment, and quality-control rules. Between-device and within-device reproducibility may both matter.

Mode and language can affect score consistency

Paper, web, app, telephone, interview, and proxy administration can differ in privacy, layout, assistance, and interpretation. Translation can change item meaning or response use. Reliability evidence should match the mode and language used in the study.

Mode-equivalence studies should examine agreement and systematic differences, not only correlation. Accessibility adaptations should be evaluated without assuming they invalidate the instrument.

Restricted range can lower reliability

A homogeneous sample has less between-person variance, which can lower a relative reliability coefficient even when measurement error is unchanged. Conversely, a very heterogeneous sample can produce an impressive coefficient that overstates precision within clinically important subgroups. Score distributions and subgroup results should accompany the overall estimate.

Floor and ceiling effects also restrict observable variation and can reduce reliability or responsiveness at the extremes. Conditional precision is useful when decisions occur near a threshold.

Sample size should support stable estimates

Reliability coefficients have sampling uncertainty. Small studies can produce wide confidence intervals and unstable variance components, especially with few raters or categories. Sample-size planning should target the expected coefficient, minimum acceptable value, confidence-interval width, number of raters, repeated measures, and attrition.

Reporting only a point estimate can make weak evidence appear definitive. Confidence intervals should be calculated with methods appropriate to the coefficient and design.

Missing repeated measurements can bias assessment

Participants who complete both assessments may be healthier, more stable, or more able to use the instrument. Excluding incomplete pairs can therefore alter score variability and reliability. Missingness by severity, group, rater, mode, and reason should be reported.

Imputation is not automatically appropriate because reliability depends on observed variation and error. The method should preserve the repeated-measure structure and be tested through sensitivity analysis.

Reliability thresholds depend on consequences

Rules of thumb can support planning but should not replace context. Group-level research may tolerate more measurement error than individual diagnosis, eligibility, or treatment decisions. The acceptable coefficient also depends on alternative measures, feasibility, and the cost of false decisions.

A pre-specified criterion should include both relative reliability and absolute agreement. A coefficient above a threshold cannot compensate for unacceptable limits of agreement.

Economic models inherit measurement reliability

Unreliable clinical measures can misclassify response, disease state, progression, or subgroup eligibility in an economic model. They can attenuate treatment effects, alter transition probabilities, and create uncertainty in mapped utilities. Models should represent the measurement rule used in the source evidence.

Relevant consequences include:

  • Misclassification around treatment-response thresholds.
  • Regression dilution in risk equations.
  • Uncertain health-state assignment.
  • Attenuated relationships used for utility mapping.
  • Different reliability across trials and routine practice.
  • Extra monitoring or confirmation required for individual decisions.

Worked reliability example

Suppose a scale has standard deviation 10 and test-retest reliability 0.84 in stable patients. The standard error of measurement is 4 points, and the 95% smallest detectable change is approximately 11.1 points.

$$ SEM=10\sqrt{1-0.84}=4 $$

$$ SDC_{95}=1.96\sqrt{2}\times4\approx11.1 $$

An individual's observed improvement of 6 points is smaller than this detectable-change threshold under the stated assumptions. That does not prove no real improvement; it means the instrument cannot distinguish the change from measurement error with the specified individual-level confidence.

Common mistakes

Reliability is often reduced to one coefficient reported without its model or population. That can conceal whether the statistic answers the intended question. The following errors should be checked before a reliability claim supports measurement or decision making.

  • Treating reliability as a fixed property of an instrument.
  • Reporting Cronbach's alpha without evidence of unidimensionality.
  • Assuming a high alpha proves validity or responsiveness.
  • Reporting an ICC without its model, type, and agreement definition.
  • Using correlation as evidence of agreement.
  • Interpreting kappa without prevalence, marginals, and raw agreement.
  • Calling the smallest detectable change a meaningful patient change.
  • Using a sample with restricted range and generalising the coefficient broadly.
  • Mixing SD and reliability estimates from incompatible populations.
  • Ignoring rater, device, language, mode, and version effects.

Reporting a reliability assessment

Transparent reporting should allow the coefficient and its interpretation to be reproduced. The report should describe the instrument, sample, replication conditions, statistical model, and absolute measurement error. Point estimates should be accompanied by uncertainty and decision-specific criteria.

  • Identify the instrument, version, domain, score, language, and mode.
  • Describe the population, setting, severity range, sample size, and missingness.
  • State the reliability form and why it fits the intended use.
  • Describe raters, occasions, interval, stability criteria, training, and blinding.
  • Report alpha or omega with structural evidence when assessing internal consistency.
  • Report the complete ICC model or kappa weighting and all assumptions.
  • Report raw agreement, SEM, limits of agreement, detectable change, and confidence intervals.
  • State the pre-specified adequacy criterion, limitations, and supported uses.

The decision standard

A reliability assessment is credible when it shows that measurement error is acceptably small relative to the distinctions the instrument must make in the intended population and setting. No single coefficient establishes that conclusion. Strong evidence combines an appropriate reliability model, absolute agreement, transparent replication conditions, and a threshold tied to whether the score will compare groups, monitor individuals, or determine care.

Library

Publications

1
  • Journal article

    GRADE Guidelines: 1. Introduction — GRADE Evidence Profiles and Summary of Findings Tables — Guyatt, Oxman, Akl, Kunz, Vist, Brozek, et al., GRADE Series ed., 2011 (Journal of Clinical Epidemiology)

    The introductory paper of the GRADE (Grading of Recommendations Assessment, Development and Evaluation) series, setting out how to rate certainty of evidence (high/moderate/low/very low) and build evidence profiles and summary-of-findings tables.

Frequently Asked Questions (6)

  • What is reliability assessment?

    An evaluation of the consistency and reproducibility of an instrument's results, covering internal consistency, inter-rater, and test-retest reliability.

    Source: Landis & Koch 1977

  • What property of an instrument does reliability assessment establish?

    Reliability assessment establishes whether an instrument produces consistent, reproducible results, so that the same thing measured under the same conditions gives much the same score each time. It draws together several checks, including agreement between raters, stability across repeated occasions, and coherence among items, each addressing a different source of inconsistency. An instrument that fails these cannot be trusted, since its scores would reflect noise as much as signal. Consistency is the property it confirms. Streiner and Norman (2008) describe this process.

    Source: Streiner & Norman 2008

  • What forms of reliability does reliability assessment cover?

    Reliability assessment covers internal consistency reliability, the coherence of an instrument's items measuring the same construct, quantified by measures such as Cronbach's alpha; inter-rater reliability, the agreement between different raters assessing the same phenomenon, quantified by kappa or the intraclass correlation coefficient; intra-rater reliability, the consistency of a single rater over occasions; and test-retest reliability, the agreement between scores from the same instrument on separate occasions when no change is expected. These forms address different sources of inconsistency. So reliability assessment examines consistency across items, raters, and time to characterise an instrument's dependability.

    Source: Cronbach 1951

  • Why is reliability assessment important?

    Reliability assessment is important because an instrument's measurements are only useful if they are consistent and reproducible; if results vary arbitrarily across items, raters, or occasions, they cannot be trusted or compared. Reliability is a prerequisite for validity, since an unreliable instrument cannot validly measure a construct. Assessing reliability ensures that the instrument gives dependable measurements suitable for research and practice. So reliability assessment matters for establishing that an instrument produces stable, consistent results, which is necessary before its scores can be relied upon and before its validity can be meaningfully established.

    Source: Landis & Koch 1977

  • How is reliability assessed?

    Reliability is assessed using statistics appropriate to each form: internal consistency by Cronbach's alpha, from the intercorrelations among items in a single administration; inter-rater and intra-rater reliability by kappa for categorical ratings or the intraclass correlation coefficient for continuous measures, comparing raters or repeated ratings; and test-retest reliability by correlating scores from the same instrument on two occasions when no real change is expected. Higher values indicate greater reliability. So reliability is assessed by quantifying the consistency of measurements across items, raters, or occasions using these statistics, evaluating the instrument's reproducibility.

    Source: Landis & Koch 1977

  • How does reliability relate to validity?

    Reliability relates to validity as a necessary but not sufficient condition: an instrument must produce consistent, reproducible results to be valid, since an unreliable measure cannot accurately capture a construct, but reliability alone does not guarantee validity, as an instrument can consistently measure the wrong thing. So reliability concerns consistency and validity concerns accuracy in measuring the intended construct. Both are needed for a good instrument, with reliability assessed first to establish consistency and validity then establishing that the consistent measurements reflect the intended concept. So reliability underpins but does not ensure validity.

    Source: Landis & Koch 1977

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 22 Sep 2026

Content version: 1.0.0

Canonical Identity

Term code
HE-ES-EA-035

Stable URI · Machine-readable · Resolvable · CC BY 4.0