Concept Architecture
Concept
Theoretically, Criterion Validity is the extent to which the measurements obtained from an instrument agree with those from an established reference standard or criterion that accurately measures the same construct. It is founded on classical test theory and psychometric validation, providing evidence that a new instrument produces results consistent with an accepted gold standard. Criterion validity exists to determine whether a measurement instrument can accurately predict or reproduce established measurements.
Mathematically, criterion validity is evaluated by quantifying the agreement or association between the instrument under evaluation and the reference criterion. Correlation coefficients, regression analyses, receiver operating characteristic (ROC) analysis and diagnostic accuracy statistics are commonly used to assess criterion validity. The strength of criterion validity is determined by the degree of correspondence between the two measures.
In practice, criterion validity is assessed by administering both the new instrument and the reference standard to the same population and comparing the resulting measurements. Health economists evaluate criterion validity when validating clinical outcome assessments, utility instruments and patient-reported outcome measures to ensure that new measures produce results consistent with established standards.
Purpose
Used to determine whether a measurement instrument produces results consistent with an accepted reference standard, supporting the validity of health outcome measures used in clinical research and economic evaluation.
Mathematical Formulae
Primary Formula
There is no universally recognised canonical mathematical formula.
Supporting Formulae
Pearson correlation coefficient:
r = Cov(X,Y) / (�?�?)
Linear regression:
Y = ?? + ??X + �
Area Under the ROC Curve:
AUC = ? TPR(FPR) d(FPR)
Related Mathematical Methods
Pearson Correlation
Linear Regression
Receiver Operating Characteristic Analysis
Area Under the Curve
Sensitivity
Specificity
Construct Validity
Example
A new electronic patient-reported quality-of-life questionnaire is compared with an established paper-based reference instrument. The scores demonstrate a Pearson correlation coefficient of r = 0.91 and excellent agreement across clinically relevant score ranges, providing strong evidence of criterion validity.
Excel Implementation
| Function | Example Formula | Health Economics Application |
|---|---|---|
| CORREL | =CORREL(B2:B101,C2:C101) | Assess agreement with the reference standard |
| RSQ | =RSQ(B2:B101,C2:C101) | Quantify shared variance between instruments |
| LINEST | =LINEST(Y_range,X_range,TRUE,TRUE) | Estimate regression relationships with the reference measure |
| FORECAST.LINEAR | =FORECAST.LINEAR(A102,C2:C101,B2:B101) | Predict criterion values from the new instrument |
| STDEV.S | =STDEV.S(B2:B101) | Assess variability during validation |
VBA (Optional)
VBA can automate criterion validity analyses by calculating agreement statistics, regression models and validation reports across multiple measurement instruments.
Sources
- Cronbach LJ, Meehl PE. Construct Validity in Psychological Tests.
- Streiner DL, Norman GR, Cairney J. Health Measurement Scales: A Practical Guide to Their Development and Use.
- DeVellis RF. Scale Development: Theory and Applications.
- COSMIN Initiative. COSMIN Methodology for Evaluating Measurement Properties.
- Portney LG, Watkins MP. Foundations of Clinical Research: Applications to Practice.
Related Concepts (2)
Library
Publications
1
Cochrane Handbook for Systematic Reviews of Interventions — Higgins, Thomas, Chandler, Cumpston, Li, Page & Welch, 2nd Edition ed., 2019 (John Wiley & Sons / Cochrane)
The standard guide to planning, conducting, interpreting and reporting systematic reviews of health interventions, with extensive material on meta-analysis, network meta-analysis, risk of bias, GRADE, equity, complex interventions and economics evidence. Maintained as a living online resource.
BookView source →
Frequently Asked Questions (6)
What is criterion validity?
The degree to which an instrument's scores correlate with an accepted external reference measure, either concurrently or in predicting a future outcome.
Source: Cronbach & Meehl 1955
What reference does criterion validity compare an instrument against?
Criterion validity compares an instrument's scores against an accepted external standard taken to reflect the truth, asking how closely the two agree. The reference might be a definitive diagnostic test measured at the same time, or a future outcome the instrument is meant to predict. Close agreement with this yardstick shows the instrument can stand in for it, which is useful when the criterion is costly, invasive, or only available later. It is validity judged against a trusted benchmark. Streiner and Norman (2008) describe this form.
Source: Streiner & Norman 2008
What are the types of criterion validity?
Criterion validity has two types: concurrent validity, in which the instrument's scores are compared with a criterion measure taken at the same time, assessing whether they agree with the established measure; and predictive validity, in which the scores are compared with a future outcome, assessing whether they predict it. Both compare the instrument against an external reference, differing in the timing of the criterion. Concurrent validity addresses agreement with a current standard, and predictive validity addresses forecasting a later outcome. So criterion validity encompasses these two forms, both relying on comparison with an accepted criterion.
Source: Cronbach & Meehl 1955
How is criterion validity assessed?
Criterion validity is assessed by correlating the instrument's scores with the criterion measure, examining whether they agree, for concurrent validity, or whether the scores predict the future outcome, for predictive validity. A strong correlation or predictive relationship supports criterion validity. The criterion must be an accepted, valid reference for the concept. Statistical measures of association or predictive accuracy quantify the relationship. So criterion validity is demonstrated by showing that the instrument's scores correspond to the criterion, either concurrently or predictively, providing direct evidence that the instrument measures the concept the criterion represents.
Source: Cronbach & Meehl 1955
When is criterion validity used?
Criterion validity is used when an accepted external reference measure, or gold standard, exists for the concept, against which the instrument's scores can be compared, either concurrently or in predicting an outcome. It provides direct validity evidence in such cases. Where no suitable criterion exists, as for many abstract concepts, construct validity is used instead, relying on theoretical expectations. So criterion validity applies when a valid criterion is available, and its use depends on the existence of an accepted reference measure, making it common for measures that can be checked against a definitive standard or a clear outcome.
Source: Cronbach & Meehl 1955
What are the limitations of criterion validity?
The limitations of criterion validity include the requirement for an accepted, valid criterion, which may not exist for abstract concepts, restricting its use; the dependence of the assessment on the quality of the criterion, since a flawed criterion undermines the comparison; and, for predictive validity, the need for follow-up to observe the outcome. If the criterion is itself imperfect, agreement with it is not conclusive. So criterion validity is applicable only when a suitable criterion is available and is limited by that criterion's validity, with construct validity used where no adequate criterion exists.
Source: Cronbach & Meehl 1955
Trust Record
Verified by Dr Darrin Baines
British health economist
Professional identity: darrinbaines.org
Verification date: 25 Nov 2025
Content version: 1.0.0
Canonical Identity
- Persistent URI
- https://healtheconomics.wiki/concept/criterion-validity
- Term code
- HE-ES-EA-010
Stable URI · Machine-readable · Resolvable · CC BY 4.0