Mean cost of an algorithm-flagged cohort weighted by the positive predictive value

When cost depends only on true disease status, the flagged cohort is a mix of true cases, in proportion PPV, and false positives, in proportion 1 minus PPV, so its mean cost is the PPV-weighted average of the two group costs. The bias against the true mean cost of cases is minus (1 minus PPV) times the cost gap. The PPV to use is the one in the population the cohort is drawn from (HE-FM-DXA-003), not the validation sample's (HE-FM-DXA-001).

Signature

C_flag = PPV * C_1 + (1 - PPV) * C_0; Bias = C_flag - C_1
Inputs
InputsDefinitionUnit
PPVShare of flagged patients who truly have the condition in the database usedproportion
C_1Mean cost of true cases, the quantity the analyst wantspounds per patient per year
C_0Mean cost of flagged patients without the condition, taken in the article as that of the average person without itpounds per patient per year
Output
C_flagMean cost per flagged patient, true cases and false positives togetherpounds per patient per year
BiasFlagged-cohort mean minus the true mean cost of casespounds per patient per year

Function

Correcting prevalence and cost estimates for misclassification by a measure validated against a criterion

Uses the sensitivity and specificity of a measure, estimated against a criterion such as medical record review, to recover the true prevalence from the share the measure flags and to show how false positives shift the mean cost of a flagged cohort. Sensitivity, specificity and predictive values themselves are HE-FM-DXA-001 (from the two-by-two table) and HE-FM-DXA-003 (at a given prevalence) on the Diagnostic Accuracy page. Notation follows the Criterion Validity article and its heart failure claims algorithm.

Computational function

  • Computational function: validation table and flagged share to corrected prevalence, database PPV and flagged-cohort cost

    Runs the whole chain from a validation study to a corrected cost estimate: sensitivity, specificity and PPV from the validation two-by-two table, the Rogan-Gladen prevalence in the database from its flagged share, the PPV there, the flagged-cohort mean cost in each population, the inflation of the cohort size and the total cost of flagged patients against true cases. The inputs differ from the formula's: the four validation counts, the database's flagged share and size, and the two group costs.

    Inputs and outputs: TP, FN, TN, FP: Validation counts against the criterion. Unit: patients.; a: Share of the database flagged. Unit: proportion.; C1, C0: Mean annual cost of true cases and of non-cases. Unit: pounds per patient.; N: Database size. Unit: patients.; se, sp, ppv_valid: Accuracy and PPV in the validation sample (HE-FM-DXA-001). Unit: proportion.; prev, ppv_db: Rogan-Gladen prevalence and PPV in the database (HE-FM-CRV-001, HE-FM-DXA-003). Unit: proportion.; cost_valid, cost_db: Flagged-cohort mean cost in each population. Unit: pounds per patient.; flagged_per_case: Flagged patients per true case, a / prev. Unit: ratio.; total_flagged, total_cases: Total annual cost of flagged patients and of true cases. Unit: pounds per year.

    Assumption: Sensitivity and specificity transfer unchanged from the validation sample to the database, and cost depends only on true status, as in the article.

    Worked example (Heart failure algorithm, article example): With 180, 20, 760 and 40 in the validation table, 6.7 per cent flagged in a database of 100,000 and costs of 6,000 and 1,000 pounds, sensitivity is 0.90, specificity 0.95, the validation PPV 0.8182 and the database prevalence 0.02, so the database PPV is 0.2687; the flagged mean cost falls from 5,090.91 to 2,343.28 pounds, the cohort is 3.35 times the true number of cases and its total cost is 15.7 million pounds against 12 million for the true cases, as in the article. TP = 180; FN = 20; TN = 760; FP = 40; a = 0.067; C1 = 6000; C0 = 1000; N = 100000; prev = 0.02; ppv_db = 0.2687; cost_db = 2343.28; flagged_per_case = 3.35

    Worked example (Database flagged at the validation rate): If the database flagged 22 per cent, the corrected prevalence would be 0.20, the same as the validation sample, and the database PPV and mean cost would equal the validation values, 0.8182 and 5,090.91 pounds (computed here for illustration). a = 0.22; prev = 0.2; ppv_db = 0.8182; cost_db = 5090.91

    Excel: With the counts named TruePos, FalseNeg, TrueNeg and FalsePos, =TruePos/(TruePos+FalseNeg) and =TrueNeg/(TrueNeg+FalsePos) return AlgSens and AlgSpec, the formula of HE-FM-CRV-001 returns TruePrev from FlagShare, =AlgSens*TruePrev/(AlgSens*TruePrev+(1-AlgSpec)*(1-TruePrev)) returns the database PPV into PPVFlag, and the formulas of HE-FM-CRV-002 return FlagMeanCost and FlagCostBias.

    R: cv_cost <- function(TP, FN, TN, FP, a, C1, C0, N) { se <- TP/(TP+FN); sp <- TN/(TN+FP); ppv_v <- TP/(TP+FP); p <- min(1, max(0, (a+sp-1)/(se+sp-1))); ppv_d <- se*p/(se*p+(1-sp)*(1-p)); cost_d <- ppv_d*C1+(1-ppv_d)*C0; list(se = se, sp = sp, ppv_valid = ppv_v, prev = p, ppv_db = ppv_d, cost_valid = ppv_v*C1+(1-ppv_v)*C0, cost_db = cost_d, flagged_per_case = a/p, total_flagged = N*a*cost_d, total_cases = N*p*C1) } Base R only; cv_cost(180, 20, 760, 40, 0.067, 6000, 1000, 100000) returns the first example.

    Python: def cv_cost(TP, FN, TN, FP, a, C1, C0, N): se = TP/(TP+FN); sp = TN/(TN+FP); ppv_v = TP/(TP+FP); p = min(1.0, max(0.0, (a+sp-1)/(se+sp-1))); ppv_d = se*p/(se*p+(1-sp)*(1-p)); cost_d = ppv_d*C1+(1-ppv_d)*C0; return {"se": se, "sp": sp, "ppv_valid": ppv_v, "prev": p, "ppv_db": ppv_d, "cost_valid": ppv_v*C1+(1-ppv_v)*C0, "cost_db": cost_d, "flagged_per_case": a/p if p > 0 else float("inf"), "total_flagged": N*a*cost_d, "total_cases": N*p*C1} Returns the same values as the R function.

    Test (Flagged patients times PPV equal the true cases found): When TruePrev is not truncated, FlagShare times PPVFlag equals AlgSens times TruePrev, the share of the database that is both flagged and a true case. Expected result: TRUE. FALSE shows the validation PPV (0.8182) used for the database, which gives 0.0548 against 0.018. Excel check: =ABS(FlagShare*PPVFlag-AlgSens*TruePrev)<1E-9

    Common error (Correcting prevalence but costing with the validation PPV): The corrected prevalence is only half the job; the cost of the flagged cohort must use the PPV at that prevalence, or the 5,091 pound validation mean is carried into a database where the flagged mean is 2,343.

    Source: Rogan WJ, Gladen B. Estimating prevalence from the results of a screening test. American Journal of Epidemiology. 1978;107(1):71-76. doi:10.1093/oxfordjournals.aje.a112510 (abstract read). Abstract; Flor M, Weiß M, Selhorst T, Müller-Graf C, Greiner M. Comparison of Bayesian and frequentist methods for prevalence estimation under misclassification. BMC Public Health. 2020;20:1135. doi:10.1186/s12889-020-09177-4 (full text read). Background, equations 1 to 3; Macaskill P, Gatsonis C, Deeks JJ, Harbord RM, Takwoingi Y. Chapter 10: Analysing and presenting results. In: Deeks JJ, Bossuyt PM, Gatsonis C, editors. Cochrane Handbook for Systematic Reviews of Diagnostic Test Accuracy. Version 1.0. The Cochrane Collaboration; 2010 (full text read). Section 10.2.3.2.

    se = TP / (TP + FN); sp = TN / (TN + FP); prev = min(1, max(0, (a + sp - 1) / (se + sp - 1))); ppv_db = se * prev / (se * prev + (1 - sp) * (1 - prev)); cost_db = ppv_db * C1 + (1 - ppv_db) * C0

Try this function

Implementations

  • Excel

    Flagged-cohort mean cost and its bias from named cells

    With PPVFlag, CostCase and CostNonCase named, the formulas return the mean cost of the flagged cohort and its bias against the cost of true cases, held in FlagMeanCost and FlagCostBias.

    =PPVFlag*CostCase+(1-PPVFlag)*CostNonCase; =FlagMeanCost-CostCase

Assumptions

  • Cost depends only on true disease status

    A flagged patient costs C_1 if a true case and C_0 otherwise, whatever the flag; if false positives are sicker than the average non-case, C_0 should be their own mean cost.

  • Predictive value of the population costed

    PPV is the share of true cases among the flagged patients actually costed, which depends on the prevalence there; HE-FM-DXA-003 computes it from sensitivity, specificity and prevalence.

Worked examples

  • Flagged cohort in the validation sample

    With 180 of 220 flagged patients true cases (PPV 0.8181818), cases at 6,000 pounds and non-cases at 1,000, the mean is about 5,091 pounds, 909 pounds or 15.2 per cent below the true 6,000, as in the article.

    PPV = 0.8181818; C_1 = 6000; C_0 = 1000; C_flag = 5090.91; Bias = -909.09
  • Flagged cohort in the full database at 2 per cent prevalence

    At the database's PPV of 0.2686567 the mean falls to about 2,343 pounds, less than half the true 6,000, as in the article.

    PPV = 0.2686567; C_1 = 6000; C_0 = 1000; C_flag = 2343.28; Bias = -3656.72
  • Flagged-cohort mean cost of an algorithm with no false positives

    With a PPV of 1 the flagged mean equals the cost of cases and the bias is zero.

    PPV = 1; C_1 = 6000; C_0 = 1000; C_flag = 6000; Bias = 0
  • Algorithm with no false positives

    With a PPV of 1 the flagged mean equals the cost of cases and the bias is zero.

    PPV = 1; C_1 = 6000; C_0 = 1000; C_flag = 6000; Bias = 0

Common errors

  • Carrying the validation sample's predictive value into routine data

    Benchimol and colleagues found that in 65.9 per cent of validation studies reporting predictive values the validation cohort had a higher prevalence than the administrative data, potentially inflating those values; in the article's example the PPV falls from 0.82 to 0.27 and the flagged mean cost from 5,091 to 2,343 pounds.

  • Correcting the mean but not the number of patients

    The flagged cohort is 3.35 times the true number of cases, so its total cost of 15.7 million pounds per 100,000 people is about 31 per cent above the 12 million spent on true cases, even though its mean cost is too low; both the count and the mean need correcting.

  • Treating the cost bias as a fixed property of the algorithm

    With sensitivity and specificity unchanged, the bias is minus 909 pounds in the validation sample and minus 3,657 pounds in the database, because PPV moves with prevalence; it has to be recomputed for each population costed.

Sources

  • Predictive values defined from the two-by-two table in the Cochrane DTA Handbook

    Macaskill P, Gatsonis C, Deeks JJ, Harbord RM, Takwoingi Y. Chapter 10: Analysing and presenting results. In: Deeks JJ, Bossuyt PM, Gatsonis C, editors. Cochrane Handbook for Systematic Reviews of Diagnostic Test Accuracy. Version 1.0. The Cochrane Collaboration; 2010 (full text read). Section 10.2.3.2: the positive predictive value is the probability that a case with a positive index test result is diseased, P(D+|T+), estimated from the table as a/(a+b); section 10.4.2: predictive values depend on prevalence.

    View source →

  • Prevalence in validation cohorts and misclassification in administrative data

    Benchimol EI, Manuel DG, To T, Griffiths AM, Rabeneck L, Guttmann A. Development and use of reporting guidelines for assessing the quality of validation studies of health administrative data. Journal of Clinical Epidemiology. 2011;64(8):821-829. doi:10.1016/j.jclinepi.2010.10.006 (abstract read). Abstract: in 271 included articles few reported four or more statistical estimates of accuracy (36.9 per cent); in 65.9 per cent of studies reporting PPV or NPV, the prevalence of disease in the validation cohort was higher than in the administrative data, potentially falsely elevating predictive values; such omissions could lead to misclassification bias and incorrect estimation of incidence and health services utilisation rates.

    View source →

Canonical Identity