Functions & Formulae

Each applied formula has its own function page, with a signature, implementations, and tests.

Correcting prevalence and cost estimates for misclassification by a measure validated against a criterion

p = (a + Sp - 1) / (Se + Sp - 1); C_flag = PPV * C_1 + (1 - PPV) * C_0

Uses the sensitivity and specificity of a measure, estimated against a criterion such as medical record review, to recover the true prevalence from the share the measure flags and to show how false positives shift the mean cost of a flagged cohort. Sensitivity, specificity and predictive values themselves are HE-FM-DXA-001 (from the two-by-two table) and HE-FM-DXA-003 (at a given prevalence) on the Diagnostic Accuracy page. Notation follows the Criterion Validity article and its heart failure claims algorithm.

  • Rogan-Gladen corrected prevalence from the share flagged by an imperfect algorithm

    p = min(1, max(0, (a + Sp - 1) / (Se + Sp - 1)))

    The share flagged, the apparent prevalence, mixes true positives and false positives: a = Se p + (1 minus Sp)(1 minus p). Solving for p gives the Rogan-Gladen estimator, which removes the false positive rate from the flagged share and divides by Youden's index, Se + Sp minus 1. Values outside the range 0 to 1 are truncated, the usual convention, and arise when the flagged share lies below the false positive rate or above the sensitivity.

  • Mean cost of an algorithm-flagged cohort weighted by the positive predictive value

    C_flag = PPV * C_1 + (1 - PPV) * C_0; Bias = C_flag - C_1

    When cost depends only on true disease status, the flagged cohort is a mix of true cases, in proportion PPV, and false positives, in proportion 1 minus PPV, so its mean cost is the PPV-weighted average of the two group costs. The bias against the true mean cost of cases is minus (1 minus PPV) times the cost gap. The PPV to use is the one in the population the cohort is drawn from (HE-FM-DXA-003), not the validation sample's (HE-FM-DXA-001).