Rogan-Gladen corrected prevalence from the share flagged by an imperfect algorithm

The share flagged, the apparent prevalence, mixes true positives and false positives: a = Se p + (1 minus Sp)(1 minus p). Solving for p gives the Rogan-Gladen estimator, which removes the false positive rate from the flagged share and divides by Youden's index, Se + Sp minus 1. Values outside the range 0 to 1 are truncated, the usual convention, and arise when the flagged share lies below the false positive rate or above the sensitivity.

Signature

p = min(1, max(0, (a + Sp - 1) / (Se + Sp - 1)))
Inputs
InputsDefinitionUnit
aShare of the database or population that the algorithm flags as having the conditionproportion
SpShare of criterion-negative people whom the algorithm does not flag, from the validation study (HE-FM-DXA-001)proportion
SeShare of criterion-positive people whom the algorithm flags, from the validation study (HE-FM-DXA-001)proportion
Output
pEstimated share of the population with the condition, truncated to the range 0 to 1proportion

Function

Correcting prevalence and cost estimates for misclassification by a measure validated against a criterion

Uses the sensitivity and specificity of a measure, estimated against a criterion such as medical record review, to recover the true prevalence from the share the measure flags and to show how false positives shift the mean cost of a flagged cohort. Sensitivity, specificity and predictive values themselves are HE-FM-DXA-001 (from the two-by-two table) and HE-FM-DXA-003 (at a given prevalence) on the Diagnostic Accuracy page. Notation follows the Criterion Validity article and its heart failure claims algorithm.

Try this function

Implementations

  • Excel

    Corrected prevalence from the flagged share and named accuracy

    With FlagShare, AlgSens and AlgSpec named, the formula returns the truncated Rogan-Gladen prevalence, held in TruePrev.

    =MIN(1,MAX(0,(FlagShare+AlgSpec-1)/(AlgSens+AlgSpec-1)))

Assumptions

  • Sensitivity and specificity transfer from the validation sample

    The correction needs the algorithm's sensitivity and specificity in the population being corrected to be the values estimated in the validation study, known and constant; the article flags this transfer as something to check.

  • Informative algorithm with Se + Sp above 1

    Youden's index must be positive; when Se + Sp equals 1 the flagged share carries no information about prevalence and the formula is undefined.

Worked examples

  • Heart failure algorithm flagging 6.7 per cent of a claims database

    With sensitivity 0.90 and specificity 0.95, (0.067 minus 0.05) / 0.85 gives a true prevalence of 0.02, as in the article.

    a = 0.067; Se = 0.9; Sp = 0.95; p = 0.02
  • Same algorithm in its validation sample

    In the validation sample the algorithm flags 220 of 1,000 patients; (0.22 minus 0.05) / 0.85 gives 0.20, the 200 of 1,000 with heart failure in the records.

    a = 0.22; Se = 0.9; Sp = 0.95; p = 0.2
  • Heart failure algorithm corrected in its validation sample

    In the validation sample the algorithm flags 220 of 1,000 patients; (0.22 minus 0.05) / 0.85 gives 0.20, the 200 of 1,000 with heart failure in the records.

    a = 0.22; Se = 0.9; Sp = 0.95; p = 0.2
  • Flagged share below the false positive rate

    If the algorithm flagged only 4 per cent, the uncorrected estimate would be (0.04 minus 0.05) / 0.85, about minus 0.0118, which is truncated to 0: the algorithm's accuracy cannot be right for that population (computed here for illustration).

    a = 0.04; Se = 0.9; Sp = 0.95; p = 0

Common errors

  • Using the flagged share as the prevalence

    The algorithm flags 6.7 per cent of the database against a true 2 per cent, so counting flagged patients as cases inflates the cohort 3.35 times even though sensitivity and specificity look high.

  • Correcting with accuracy estimated in another population

    Lewis and Torgerson note that the Rogan-Gladen estimator requires sensitivity and specificity to be known and constant, and that estimates above one or below zero mean the values are wrong for the population at hand; accuracy from an enriched validation sample may not hold in routine data.

  • Reporting a negative corrected prevalence

    Flor and colleagues note that the estimate is a valid proportion only if the false positive rate is below sensitivity and the flagged share lies between them, and that it is truncated to the range 0 to 1 in practice; a negative value signals a mismatch, not a usable estimate.

Sources

  • Rogan and Gladen on estimating prevalence from a screening test

    Rogan WJ, Gladen B. Estimating prevalence from the results of a screening test. American Journal of Epidemiology. 1978;107(1):71-76. doi:10.1093/oxfordjournals.aje.a112510 (abstract read). Abstract: a common use of tests is to estimate the prevalence of disease; the frequency of positive tests is shown to be a bad estimate, and the necessary adjustments are given.

    View source →

  • Rogan-Gladen estimator, apparent prevalence and truncation

    Flor M, Weiß M, Selhorst T, Müller-Graf C, Greiner M. Comparison of Bayesian and frequentist methods for prevalence estimation under misclassification. BMC Public Health. 2020;20:1135. doi:10.1186/s12889-020-09177-4 (full text read). Background: the apparent prevalence is AP = Se times pi + (1 minus Sp)(1 minus pi) (eq. 1); solving for the prevalence gives the Rogan-Gladen estimate (AP minus (1 minus Sp)) / (Se minus (1 minus Sp)) (eq. 2), truncated in practice to the interval from 0 to 1 (eq. 3); a valid proportion needs 1 minus Sp below Se, 1 minus Sp no greater than AP and AP no greater than Se (eqs. 4a to 4c).

    View source →

  • Rogan-Gladen estimator needs known and constant sensitivity and specificity

    Lewis FI, Torgerson PR. A tutorial in estimating the prevalence of disease in humans and animals in the absence of a gold standard diagnostic. Emerging Themes in Epidemiology. 2012;9:9. doi:10.1186/1742-7622-9-9 (full text read). Section on the simple model: the classic Rogan-Gladen estimator of true prevalence requires that Se and Sp are both known and constant, which may be unrealistic, and it can produce estimates of prevalence that exceed one or are negative, in which case Se and Sp must be incorrect for the population of interest.

    View source →

Canonical Identity