Signature
C_min = mu - 1; C_max = 1 - mu; W = C / (1 - mu)
| Inputs | Definition | Unit |
|---|---|---|
mu | Mean of the binary variable, the proportion of the population with the value 1, strictly between 0 and 1 | proportion |
C | Standard concentration index of the binary variable by fractional rank in the living standards distribution, as in HE-FM-HINQ-005 | none |
C_min | Smallest feasible concentration index, reached when everyone with the value 1 is poorer than everyone with the value 0 | none |
|---|---|---|
C_max | Largest feasible concentration index, reached when everyone with the value 1 is richer than everyone with the value 0 | none |
W | Concentration index divided by 1 minus the mean of the binary variable, ranging from minus 1 to 1 | none |
Function
Bound corrections of the concentration index for binary and bounded health variables
Maps the standard concentration index of a health variable, its mean and the bounds of its measurement scale to an index whose feasible range does not depend on the mean. For a binary variable the standard index can only lie between mu minus 1 and 1 minus mu in large samples, so Wagstaff divides it by 1 minus the mean, while Erreygers multiplies it by 4 times the mean over the range of the scale, which equals 8 times the covariance of health and fractional rank over that range. The standard index itself, C = 2 cov_hr / mu, is HE-FM-HINQ-005, computed from grouped data in HE-CF-HINQ-001; the curve form by trapezoids is HE-FM-VEQ-002, the horizontal inequity index C_M minus C_N is HE-FM-HEQ-003 and the rule of 75 is HE-FM-HEQ-004. Notation follows the Concentration Index article.
Computational function
Computational function: standard, Wagstaff and Erreygers concentration indices from group values, shares and scale bounds
Takes the mean of a bounded health variable in each socioeconomic group, ordered from poorest to richest, the share of the population in each group and the bounds of the measurement scale, and returns the standard concentration index with its generalised Wagstaff and Erreygers corrections, plus the same three indices for the mirror variable. It builds each group's fractional rank as the cumulative share before the group plus half its own share, then the population-weighted mean and covariance, and then applies HE-FM-HINQ-005, HE-FM-CIX-001 in the generalised form of Erreygers and Van Ourti, and HE-FM-CIX-002. The inputs therefore differ from the formulas' variables, which take C and mu as given. Individual records can be entered one per row with their survey weights as the shares.
Inputs and outputs:
h_j: Mean of the health variable in group j, from the poorest group to the richest, or one value per person; required, between a and b. Unit: units of the health variable.;f_j: Population share or survey weight of group j, in the same order; required, above zero, rescaled to sum to 1. Unit: proportion.;a: Lower bound of the measurement scale, 0 for a binary variable. Unit: units of the health variable.;b: Upper bound of the measurement scale, 1 for a binary variable. Unit: units of the health variable.;r_j: Fractional rank of group j, an intermediate output. Unit: none.;mu: Population-weighted mean. Unit: units of the health variable.;cov_hr: Population-weighted covariance of health and fractional rank, an intermediate output. Unit: units of the health variable.;C: Standard concentration index. Unit: none.;W: Generalised Wagstaff index, C divided by 1 minus mu for a binary variable. Unit: none.;E: Erreygers corrected index. Unit: none.;C_s,W_s,E_s: The same three indices for the mirror variable a plus b minus h_j, such as freedom from illness when h_j is illness. Unit: none.Assumption: Groups are ranked by living standards from poorest to richest, each group carries its mean so only between-group inequality is measured, a and b are fixed by the measurement scale, and mu lies strictly between a and b for the Wagstaff index.
Worked example (Limiting illness across five income quintiles): The article's prevalences of 0.40, 0.30, 0.25, 0.15 and 0.10 with 20 per cent of the population in each quintile give ranks of 0.1 to 0.9, a mean of 0.24 and a covariance of minus 0.03, so C is minus 0.25, W about minus 0.3289 and E minus 0.24.
mu = 0.24; cov_hr = -0.03; a = 0; b = 1; C = -0.25; W = -0.3289; E = -0.24Worked example (Mirror variable, free of limiting illness): Good health of 0.60 to 0.90 has a mean of 0.76 and a covariance of 0.03, so C is about 0.0789 while W and E are exact mirrors of the illness values, equal to C_s, W_s and E_s from the illness run.
mu = 0.76; cov_hr = 0.03; a = 0; b = 1; C = 0.0789; W = 0.3289; E = 0.24Worked example (Unequal group sizes): An illustrative variation with the same prevalences and shares of 0.30, 0.25, 0.20, 0.15 and 0.10 from the poorest group to the richest gives ranks of 0.15, 0.425, 0.65, 0.825 and 0.95, a mean of 0.2775 and a covariance of minus 0.0283125.
mu = 0.2775; cov_hr = -0.0283125; a = 0; b = 1; C = -0.2041; W = -0.2824; E = -0.2265Excel:
=LET(h,Health,f,Share/SUM(Share),rk,SCAN(0,f,LAMBDA(acc,x,acc+x))-f/2,mu,SUMPRODUCT(f,h),cv,SUMPRODUCT(f,(h-mu)*(rk-0.5)),HSTACK(2*cv/mu,2*(UpperBound-LowerBound)*cv/((UpperBound-mu)*(mu-LowerBound)),8*cv/(UpperBound-LowerBound)))Excel 365; with group means in a column named Health, shares or weights in Share and the scale bounds in LowerBound and UpperBound, the formula spills C, W and E into three cells.R:
ci_bounded <- function(h, f, a = 0, b = 1) { f <- f/sum(f); r <- cumsum(f)-f/2; mu <- sum(f*h); cv <- sum(f*(h-mu)*(r-0.5)); mu_s <- a+b-mu; c(C = 2*cv/mu, W = 2*(b-a)*cv/((b-mu)*(mu-a)), E = 8*cv/(b-a), C_s = -2*cv/mu_s, W_s = -2*(b-a)*cv/((b-mu_s)*(mu_s-a)), E_s = -8*cv/(b-a)) }Base R;ci_bounded(c(0.40, 0.30, 0.25, 0.15, 0.10), rep(0.2, 5))returns C of minus 0.25, W of minus 0.3289 and E of minus 0.24.Python:
def ci_bounded(h, f, a=0, b=1): w = [x/sum(f) for x in f]; r = [sum(w[:j])+w[j]/2 for j in range(len(w))]; mu = sum(p*x for p, x in zip(w, h)); cv = sum(p*(x-mu)*(q-0.5) for p, x, q in zip(w, h, r)); mu_s = a+b-mu; return {'C': 2*cv/mu, 'W': 2*(b-a)*cv/((b-mu)*(mu-a)), 'E': 8*cv/(b-a), 'C_s': -2*cv/mu_s, 'W_s': -2*(b-a)*cv/((b-mu_s)*(mu_s-a)), 'E_s': -8*cv/(b-a)}Plain Python with no imports; returns the same values as the R function.Test (Mirror property of the corrections): For any input W_s equals minus W and E_s equals minus E, while C_s equals minus C only when mu is halfway between a and b; in the quintile example C_s is about 0.0789 against minus C of 0.25. Expected result: TRUE. Excel check:
=AND(ABS(WIll+WGood)<1E-9,ABS(EIll+EGood)<1E-9)Test (Sum form of the standard index): With shares that sum to 1, twice the weighted sum of health times rank divided by the mean, minus 1, returns the same C, minus 0.25 in the quintile example. Expected result: TRUE. Excel check:
=ABS(2*SUMPRODUCT(Share,Health,Rank)/SUMPRODUCT(Share,Health)-1-ConcIndex)<1E-9Common error (Groups entered from richest to poorest): Reversing the order reverses every rank and so every sign: the illness example returns C of 0.25 and E of 0.24, which reads as illness concentrated among the better-off.
Source: O'Donnell O, van Doorslaer E, Wagstaff A, Lindelow M. Analyzing health equity using household survey data. Washington, DC: World Bank; 2008. Chapter 8, equations 8.3 and 8.5 and the section on properties; Erreygers G, Van Ourti T. Journal of Health Economics. 2011;30(4):685-694. doi:10.1016/j.jhealeco.2011.04.004. Equations 2a, 4 and 5; Wagstaff A. Health Economics. 2005;14(4):429-432. doi:10.1002/hec.953.
r_j = sum_(k=1)^(j-1) [f_k] + f_j / 2; mu = sum_(j=1)^J [f_j * h_j]; cov_hr = sum_(j=1)^J [f_j * (h_j - mu) * (r_j - 0.5)]; C = 2 * cov_hr / mu; W = 2 * (b - a) * cov_hr / ((b - mu) * (mu - a)); E = 8 * cov_hr / (b - a); C_s = -2 * cov_hr / (a + b - mu); W_s = -W; E_s = -E
Try this function
Implementations
Excel
Wagstaff normalised concentration index and binary bounds in Excel
With the standard index in a cell named ConcIndex and the mean of the binary variable in MeanBinary, the three formulas return the lower bound, the upper bound and the normalised index, or #N/A when the mean is not strictly between 0 and 1.
=MeanBinary-1; =1-MeanBinary; =IF(OR(MeanBinary<=0,MeanBinary>=1),NA(),ConcIndex/(1-MeanBinary))
Assumptions
Binary variable and large samples for the Wagstaff bounds
The variable takes only the values 0 and 1, and mu is a proportion strictly between 0 and 1. The World Bank guide gives the bounds mu minus 1 and 1 minus mu for large samples; with the mid-rank (i minus 0.5)/N of HE-FM-HINQ-005 and a whole number of people with the value 1, the bounds are reached exactly when those people are the poorest or the richest (computed here for illustration). The bounds do not apply to a variable that is not binary.
Same sample, weights and ranking for the index and mean in the Wagstaff normalisation
C and mu come from the same sample, survey weights and living standards ranking, and the normalisation rescales the index without changing its sign. For a variable bounded by a and b, Erreygers and Van Ourti give the generalised Wagstaff index as (b minus a) times mu times C, divided by (b minus mu) times (mu minus a), which reduces to C divided by 1 minus mu when a is 0 and b is 1. Like the Erreygers index of HE-FM-CIX-002, it passes their mirror and cardinal invariance tests.
Worked examples
Wagstaff normalised index of limiting illness across five income quintiles
In the article's illustrative example the share of adults with a limiting long-term illness falls from 0.40 in the poorest quintile to 0.10 in the richest, with a mean of 0.24 and a standard index of minus 0.25. The feasible range is minus 0.76 to 0.76, and dividing by 0.76 gives a normalised index of about minus 0.3289, as in the article.
C = -0.25; mu = 0.24; C_min = -0.76; C_max = 0.76; W = -0.3289
Wagstaff normalised index of being free of limiting illness
The share without limiting illness, 0.60 in the poorest quintile rising to 0.90 in the richest, has a mean of 0.76 and a standard index of 0.06 divided by 0.76, about 0.0789. Its feasible range is only minus 0.24 to 0.24, and the normalised index is about 0.3289, the exact mirror of the illness value, as in the article.
C = 0.078947; mu = 0.76; C_min = -0.24; C_max = 0.24; W = 0.3289
Wagstaff normalised immunisation index at 50 per cent coverage
In an illustrative country where half of children are fully immunised and the standard index is 0.10, the feasible range is minus 0.5 to 0.5 and the normalised index is 0.2 (computed here for illustration).
C = 0.10; mu = 0.5; C_min = -0.5; C_max = 0.5; W = 0.2
Wagstaff normalised immunisation index at 90 per cent coverage
At 90 per cent coverage the same standard index of 0.10 is the largest value the data allow, since the feasible range is minus 0.1 to 0.1, so the normalised index is 1 and every unimmunised child is poorer than every immunised child. The two standard indices are equal while the normalised indices differ fivefold (computed here for illustration).
C = 0.10; mu = 0.9; C_min = -0.1; C_max = 0.1; W = 1
Common errors
Comparing standard concentration indices of binary outcomes with different means
In the article's example the standard index of limiting illness is minus 0.25 and that of being free of illness about 0.0789, so the same distribution looks about three times more unequal when described as illness. Comparing immunisation, mortality or other binary outcomes across countries or years with different means needs a bound correction, and the report should state which correction was used.
Entering the mean of a binary variable as a percentage in the Wagstaff normalisation
With mu entered as 24 in place of 0.24, 1 minus mu becomes minus 23 and the illness index of minus 0.25 turns into about 0.0109, close to zero and with the wrong sign. The mean and the index must both be on the proportion scale.
Applying the binary bounds to a count or continuous health variable
The bounds mu minus 1 and 1 minus mu hold only for a variable coded 0 and 1. For a mean of 2.5 outpatient visits a year, 1 minus mu is minus 1.5, which would put the upper bound below zero and reverse the sign of the normalised index. Non-negative count and continuous variables keep the standard index, and bounded scores need the generalised Wagstaff form or the Erreygers index with the bounds of the scale.
Sources
Wagstaff on the bounds of the concentration index for a binary variable
Wagstaff A. The bounds of the concentration index when the variable of interest is binary, with an application to immunization inequality. Health Economics. 2005;14(4):429-432. doi:10.1002/hec.953. Abstract: for a binary variable the minimum and maximum of the concentration index are the mean minus 1 and 1 minus the mean, so the range shrinks as the mean rises.
World Bank guide on normalising the concentration index of a binary variable
O'Donnell O, van Doorslaer E, Wagstaff A, Lindelow M. Analyzing health equity using household survey data: a guide to techniques and their implementation. Washington, DC: World Bank; 2008. Chapter 8, The concentration index, section on properties: for large samples the lower bound is the mean minus 1 and the upper bound 1 minus the mean, caution in comparing child mortality and immunisation across countries with different means, and normalisation by dividing through by 1 minus the mean.
Canonical Identity
Stable URI · Machine-readable · Resolvable · CC BY 4.0