Prevalence ratio from a cross-sectional two-by-two table

Divides the proportion of exposed people who have the condition by the proportion of unexposed people who have it, using the four cells of a cross-sectional two-by-two table with exposure in rows. A value of 1 means the condition is equally common in both groups, above 1 more common among the exposed and below 1 less common. Because P_1 cannot exceed 1, the ratio cannot exceed 1 / P_0.

Signature

P_1 = a / (a + b); P_0 = c / (c + d); PR = P_1 / P_0
Inputs
InputsDefinitionUnit
aNumber of exposed people who have the condition at the time of measurementpeople
bNumber of exposed people who do not have the condition at the time of measurementpeople
cNumber of unexposed people who have the condition at the time of measurement, above zeropeople
dNumber of unexposed people who do not have the condition at the time of measurementpeople
Output
P_1Proportion of exposed people who have the condition at the time of measurement, a / (a + b)proportion
P_0Proportion of unexposed people who have the condition at the time of measurement, c / (c + d), above zeroproportion
PRPrevalence among the exposed divided by prevalence among the unexposedratio, zero or above

Function

Prevalence ratio function for cross-sectional comparisons and attributable cost inputs

Maps the prevalence of a condition among exposed people and among unexposed people, measured at the same time, to their ratio, and carries that ratio into the prevalence odds ratio, a population attributable fraction of prevalent cases and subgroup prevalences for cost-of-illness and budget impact work. Applying the attributable fraction to aggregate annual costs is the top-down attributable burden HE-FM-COI-002, and the attack rate ratio HE-FM-ATR-003 has the same ratio form for new cases in an outbreak. Survey-weighted prevalence and the link between prevalence, incidence and duration are HE-FM-XSD-001 and HE-FM-XSD-002. Notation follows the Prevalence Ratio article.

Computational function

  • Computational function: attributable annual cost of prevalent cases from survey counts and a prevalence ratio

    Takes the four counts of a cross-sectional survey and three facts about a local population, its size, its exposed share and an annual cost per prevalent case, and returns the attributable annual cost of prevalent cases. It chains three steps: the prevalence ratio HE-FM-PR-001 from the survey, the attributable fraction HE-FM-PR-003 with the local exposed share, and the top-down burden HE-FM-COI-002 applied to the local prevalent cases. The inputs therefore differ from the formula's variables: the survey supplies the ratio and the prevalences, and the local population supplies the cases and costs.

    Inputs and outputs: a, b: Exposed survey respondents with and without the condition; required. Unit: people.; c, d: Unexposed survey respondents with and without the condition, with c above zero; required. Unit: people.; N: Size of the local population whose costs are attributed; required. Unit: people.; e: Exposed share of the local population; required. Unit: proportion.; k: Annual cost to the health service per prevalent case, the same in both groups; required. Unit: currency per case per year.; P_1, P_0: Survey prevalences among the exposed and the unexposed, returned as intermediate outputs. Unit: proportion.; PR: Survey prevalence ratio. Unit: ratio.; PAF: Attributable fraction of local prevalent cases. Unit: proportion.; M: Local prevalent cases at the survey prevalences. Unit: people.; A: Attributable prevalent cases. Unit: people.; C_A: Attributable annual cost. Unit: currency per year.

    Assumption: The survey prevalences and their ratio hold in the local population, the ratio is causal and unconfounded, and each prevalent case costs the same in both groups. The result is an annual cost of prevalent cases, not a lifetime cost of new cases.

    Worked example (The article's survey and local health economy): The survey gives a PR of 2.00; with 500,000 adults, 40% exposed and an illustrative GBP 1,200 per case, 30,000 of 105,000 prevalent cases are attributable and the annual cost is GBP 36.0 million. a = 180; b = 420; c = 210; d = 1190; N = 500000; e = 0.40; k = 1200; P_1 = 0.30; P_0 = 0.15; PR = 2.00; PAF = 0.285714; M = 105000; A = 30000; C_A = 36000000

    Worked example (Low prevalences with the same ratio): Survey prevalences of 0.02 and 0.01 give the same PR and PAF, but only 7,000 local prevalent cases, so 2,000 cases and GBP 2.4 million are attributable. a = 12; b = 588; c = 14; d = 1386; N = 500000; e = 0.40; k = 1200; P_1 = 0.02; P_0 = 0.01; PR = 2.00; PAF = 0.285714; M = 7000; A = 2000; C_A = 2400000

    Worked example (Equal prevalences): With 10% affected in both survey groups the PR is 1 and nothing is attributable, a limiting case that checks the implementation. a = 60; b = 540; c = 140; d = 1260; N = 500000; e = 0.40; k = 1200; P_1 = 0.10; P_0 = 0.10; PR = 1.00; PAF = 0; M = 50000; A = 0; C_A = 0

    Excel: =LET(pExp,ExposedCases/(ExposedCases+ExposedFree),pUnexp,UnexposedCases/(UnexposedCases+UnexposedFree),ratio,pExp/pUnexp,frac,ExposedShare*(ratio-1)/(1+ExposedShare*(ratio-1)),cases,Adults*(ExposedShare*pExp+(1-ExposedShare)*pUnexp),HSTACK(ratio,frac,cases,frac*cases,frac*cases*CostPerCase)) Excel 365; with the survey counts, Adults, ExposedShare and CostPerCase in named cells, the result spills PR, PAF, M, A and C_A into five cells.

    R: pr_attributable_cost <- function(a, b, c, d, N, e, k) { P_1 <- a/(a+b); P_0 <- c/(c+d); PR <- P_1/P_0; PAF <- e*(PR-1)/(1+e*(PR-1)); M <- N*(e*P_1+(1-e)*P_0); A <- PAF*M; list(PR = PR, PAF = PAF, M = M, A = A, C_A = A*k) } Base R; pr_attributable_cost(180, 420, 210, 1190, 500000, 0.40, 1200) returns a PR of 2, 30000 attributable cases and a cost of 3.6e+07.

    Python: def pr_attributable_cost(a, b, c, d, N, e, k): P_1 = a/(a+b); P_0 = c/(c+d); PR = P_1/P_0; PAF = e*(PR-1)/(1+e*(PR-1)); M = N*(e*P_1+(1-e)*P_0); A = PAF*M; return {"PR": PR, "PAF": PAF, "M": M, "A": A, "C_A": A*k} Plain Python with no imports; returns the same values as the R function.

    Test (Attributable cases equal the excess among the exposed): The attributable cases equal N x e x (P_1 minus P_0); with the article's figures 200,000 x 0.15 = 30,000. Expected result: TRUE, with the spilled result range named Result. Excel check: =ABS(INDEX(Result,1,4)-Adults*ExposedShare*(ExposedCases/(ExposedCases+ExposedFree)-UnexposedCases/(UnexposedCases+UnexposedFree)))<1E-6

    Test (Article's attributable cost): The article's survey and local figures return GBP 36,000,000. Expected result: TRUE. Excel check: =ROUND(INDEX(Result,1,5),0)=36000000

    Common error (Using the prevalence odds ratio from a logistic model): Feeding the survey POR of about 2.4286 into the attributable fraction raises the attributable cases to about 38,182 and the cost to GBP 45.8 million, 27% above the figure from the PR, because the condition is common.

    Source: Jo C. Cost-of-illness studies: concepts, scopes, and methods. Clinical and Molecular Hepatology. 2014;20(4):327-337. Section on top-down, bottom-up and econometric approaches (PAF from a prevalence and an unadjusted relative risk, applied to aggregate costs); Tamhane AR, Westfall AO, Burkholder GA, Cutter GR. Prevalence odds ratio versus prevalence ratio: choice comes with consequences. Statistics in Medicine. 2016;35(30):5730-5735. Table 3 (prevalence ratio from a two-by-two table).

    P_1 = a / (a + b); P_0 = c / (c + d); PR = P_1 / P_0; PAF = e * (PR - 1) / (1 + e * (PR - 1)); M = N * (e * P_1 + (1 - e) * P_0); A = PAF * M; C_A = A * k

Try this function

Implementations

  • Excel

    Prevalence ratio from the four table cells in Excel

    With the counts in named cells ExposedCases (a), ExposedFree (b), UnexposedCases (c) and UnexposedFree (d), Excel returns the PR, or #N/A when no unexposed person has the condition.

    =IF(UnexposedCases=0,NA(),(ExposedCases/(ExposedCases+ExposedFree))/(UnexposedCases/(UnexposedCases+UnexposedFree)))

Assumptions

  • Same case definition and timing in both prevalence ratio groups

    Both groups use the same case definition and the same timing, either point prevalence at one date or period prevalence over the same interval, and both prevalences come from one cross-sectional sample. A ratio of prevalences measured at different times or with different definitions mixes the exposure with the measurement.

  • Prevalence ratio reflects onset and duration together

    Prevalence carries both incidence and the duration of the condition, so a PR above 1 can arise because exposed people develop the condition more often, because they live with it for longer, or because unexposed people with the condition die or recover sooner. Reading the PR as a ratio of risks of becoming ill needs further assumptions about duration and survival.

Worked examples

  • Prevalence ratio of 2.00 in the article's health survey

    A survey of 2,000 adults finds the condition in 180 of 600 exposed adults and 210 of 1,400 unexposed adults. The prevalences are 0.30 and 0.15, so the PR is 2.00 and the prevalence difference is 15 percentage points.

    a = 180; b = 420; c = 210; d = 1190; P_1 = 0.30; P_0 = 0.15; PR = 2.00
  • Prevalence ratio when both prevalences are low

    With the condition in 12 of 600 exposed adults and 14 of 1,400 unexposed adults, the prevalences are 0.02 and 0.01 and the PR is again 2.00. HE-EX-PR-005 shows that the prevalence odds ratio is then close to the PR.

    a = 12; b = 588; c = 14; d = 1386; P_1 = 0.02; P_0 = 0.01; PR = 2.00
  • Prevalence ratio of 1 when the two prevalences are equal

    With the condition in 60 of 600 exposed adults and 140 of 1,400 unexposed adults, both prevalences are 0.10 and the PR is 1.00, a limiting case that checks the implementation.

    a = 60; b = 540; c = 140; d = 1260; P_1 = 0.10; P_0 = 0.10; PR = 1.00

Common errors

  • Inverting a prevalence ratio for the complement of the condition

    The PR for being free of the condition is (1 minus P_1) divided by (1 minus P_0), not the reciprocal of the PR. In the article's survey it is 0.70 / 0.85 = 0.82, whereas 1 / 2.00 = 0.50. A model that needs the complement computes it from the two prevalences.

  • Using a prevalence ratio as the risk ratio for new cases

    A PR used as the effect of an exposure on new cases treats differences in duration as differences in onset. An exposure that improves survival with a condition raises its prevalence without raising its incidence, so a prevention model built on the PR can misstate the cases avoided.

Sources

  • Tamhane and colleagues on the two-by-two prevalence ratio

    Tamhane AR, Westfall AO, Burkholder GA, Cutter GR. Prevalence odds ratio versus prevalence ratio: choice comes with consequences. Statistics in Medicine. 2016;35(30):5730-5735. Table 3: general two-by-two table for a cross-sectional study, with the prevalence ratio for the outcome and for its complement, which is not the reciprocal; Discussion on reciprocity.

    View source →

  • CDC on prevalence, incidence and duration behind a prevalence ratio

    Centers for Disease Control and Prevention. Principles of Epidemiology in Public Health Practice. 3rd ed. Lesson 3: Measures of Risk, section 2 (prevalence is based on both incidence and duration of illness; point and period prevalence).

    View source →

Canonical Identity

Stable URI · Machine-readable · Resolvable · CC BY 4.0