Survey-weighted prevalence from a cross-sectional health survey with three weighting classes

Each respondent counts in proportion to the survey weight, which corrects for unequal selection, non-response and the population profile, so the prevalence is the weighted share of respondents with the condition, the sum over respondents i of w_i y_i divided by the sum of w_i, with y_i equal to 1 for a respondent with the condition and 0 otherwise. When respondents in a weighting class share one weight, the sum reduces to the class form below; three classes are written out so that the calculator can run, and SUMPRODUCT in Excel takes any number. The formula gives the point estimate only: its standard error needs the strata and primary sampling units of the design.

Signature

P_w = (w_1 * d_1 + w_2 * d_2 + w_3 * d_3) / (w_1 * n_1 + w_2 * n_2 + w_3 * n_3)
Inputs
InputsDefinitionUnit
w_1Weight carried by each respondent in class 1, such as an over-sampled regionweight per respondent
d_1Number of respondents in class 1 who report or have the conditionrespondents
w_2Weight carried by each respondent in class 2weight per respondent
d_2Number of respondents in class 2 who report or have the conditionrespondents
w_3Weight carried by each respondent in class 3weight per respondent
d_3Number of respondents in class 3 who report or have the conditionrespondents
n_1Number of respondents in class 1 with a valid answerrespondents
n_2Number of respondents in class 2 with a valid answerrespondents
n_3Number of respondents in class 3 with a valid answerrespondents
Output
P_wWeighted share of respondents who have the conditionproportion

Function

Prevalence estimation from a cross-sectional survey and its relation to incidence and duration

Estimates the prevalence of a condition from a single survey round, weighting respondents for how the sample was drawn, and links point prevalence to incidence and mean duration in a steady state. The same link shows why averages taken over prevalent cases weight each form of a condition by incidence times duration, not by incidence alone. Notation follows the Cross-Sectional Design article and its prevalent-case utility example.

Try this function

Implementations

  • Excel

    Weighted prevalence from named weighting-class ranges

    With the class weights, case counts and respondent counts in ranges named ClassWeights, ClassCases and ClassSizes, in the same order, the formula returns the weighted prevalence, held in WtdPrev. For respondent-level data with a 0 or 1 condition flag, SUMPRODUCT of weights and flags over SUM of weights gives the same result.

    =SUMPRODUCT(ClassWeights,ClassCases)/SUMPRODUCT(ClassWeights,ClassSizes)

Assumptions

  • Weights supplied by the survey and used as given

    The weights are the survey's own, such as the Health Survey for England's selection, non-response and calibration weights; multiplying every weight by the same constant, for example to gross up to the population, leaves P_w unchanged.

  • One measurement occasion and valid answers only

    Each respondent is counted once, at the survey date, and the n counts exclude missing answers; if missingness depends on the condition, the weights do not correct for it.

Worked examples

  • Three weighting classes with an over-sampled region

    Classes of 1,000, 2,000 and 1,000 respondents with weights 0.6, 1.2 and 1.0 and 90, 150 and 60 cases give 294 / 4,000, a weighted prevalence of 0.0735 against an unweighted 300 / 4,000 or 0.075 (computed here for illustration).

    w_1 = 0.6; w_2 = 1.2; w_3 = 1; d_1 = 90; d_2 = 150; d_3 = 60; n_1 = 1000; n_2 = 2000; n_3 = 1000; P_w = 0.0735
  • Same survey with equal weights

    With every weight equal to 1 the estimate is the unweighted share, 0.075 (computed here for illustration).

    w_1 = 1; w_2 = 1; w_3 = 1; d_1 = 90; d_2 = 150; d_3 = 60; n_1 = 1000; n_2 = 2000; n_3 = 1000; P_w = 0.075
  • Three-class survey prevalence with equal weights

    With every weight equal to 1 the estimate is the unweighted share, 0.075 (computed here for illustration).

    w_1 = 1; w_2 = 1; w_3 = 1; d_1 = 90; d_2 = 150; d_3 = 60; n_1 = 1000; n_2 = 2000; n_3 = 1000; P_w = 0.075
  • Three-class survey prevalence with weights grossed up to the population

    Multiplying the weights by 1,000 to represent 4 million people leaves the prevalence at 0.0735 (computed here for illustration).

    w_1 = 600; w_2 = 1200; w_3 = 1000; d_1 = 90; d_2 = 150; d_3 = 60; n_1 = 1000; n_2 = 2000; n_3 = 1000; P_w = 0.0735
  • Same survey with weights grossed up to the population

    Multiplying the weights by 1,000 to represent 4 million people leaves the prevalence at 0.0735 (computed here for illustration).

    w_1 = 600; w_2 = 1200; w_3 = 1000; d_1 = 90; d_2 = 150; d_3 = 60; n_1 = 1000; n_2 = 2000; n_3 = 1000; P_w = 0.0735

Common errors

  • Analysing a complex survey as a simple random sample

    STROBE asks cross-sectional studies to describe analytical methods that take account of the sampling strategy; ignoring the weights gives 0.075 instead of 0.0735 in the illustrative example, and ignoring clustering and stratification also misstates the standard error.

  • Quoting a simple random sample standard error for a weighted estimate

    The Health Survey for England methods report notes that standard errors and confidence intervals for its estimates are generally larger than those of an unweighted simple random sample of the same size, and expresses this through the design factor, the ratio of the two standard errors.

Sources

  • Inverse-probability weighting of survey observations

    Lumley T. Analysis of complex survey samples. Journal of Statistical Software. 2004;9(8):1-19. doi:10.18637/jss.v009.i08 (full text read). Section 2.1: units sampled with unequal probability must receive correspondingly unequal weights; inverse-probability weighting is needed for valid point estimates. Section 2.2.1: the contribution of each primary sampling unit to a population total is the sum of its observations divided by their sampling probabilities, and the variance is estimated from the primary sampling units within strata. The weighted prevalence is the ratio of two such weighted totals, of cases and of respondents (step derived in this record).

    View source →

  • Weighting and design factors in the Health Survey for England 2022

    NHS England. Health Survey for England 2022: Methods. Leeds: NHS England; 2024 (full text read). Section 7: selection weights for addresses, dwelling units and households, non-response weights and calibration to ONS mid-year population estimates. Section 8.2: standard errors and confidence intervals are generally larger than those of an unweighted simple random sample of the same size; the ratio of the standard error of the complex sample to that of a simple random sample of the same size is the design factor.

    View source →

  • STROBE item 12(d) on analysis that takes account of the sampling strategy

    Vandenbroucke JP, von Elm E, Altman DG, Gøtzsche PC, Mulrow CD, Pocock SJ, et al. Strengthening the Reporting of Observational Studies in Epidemiology (STROBE): explanation and elaboration. PLoS Medicine. 2007;4(10):e297. doi:10.1371/journal.pmed.0040297 (full text read). Item 12(d): cross-sectional study, if applicable, describe analytical methods taking account of sampling strategy; explanation: measures of precision should be corrected using the design effect, a ratio measure of how much precision is gained or lost under a more complex sampling strategy than simple random sampling.

    View source →

Canonical Identity