Signature
mu = (1 / N) * sum_(i=1)^N [Delta_i * A_i / S_c_i]
| Inputs | Definition | Unit |
|---|---|---|
N | All patients, censored or not | count |
Delta_i | 1 if patient i died or was followed to the time limit, 0 if censored earlier | none |
A_i | Cumulative cost of patient i up to death, censoring or the time limit | currency |
S_c_i | Kaplan-Meier estimate, with censoring as the event, of remaining uncensored at the complete case's death or limit time; above 0 | probability |
mu | Estimated mean cost per patient up to the time limit | currency |
|---|
Function
Inverse probability of censoring weighting for dependent censoring
Maps patients who remain under observation to weights equal to the inverse of their probability of having remained uncensored, given their covariate history, so that they also represent similar patients who were censored. Weighted versions of the Kaplan-Meier estimator, the Cox model and mean cost estimators then remove the bias that dependent censoring causes, provided every factor predicting both censoring and outcome is measured. The notation follows the Dependent Censoring article and its switchers-censored control arm.
Computational function
Computational function: Bang and Tsiatis weighted mean cost from patient-level censored data
Computes the Bang and Tsiatis simple weighted mean cost from patient-level data: it sorts patients by time, estimates the Kaplan-Meier probability of remaining uncensored with censoring as the event, weights each complete case by its inverse and averages over all patients, returning also the full-sample and complete-case means for comparison. The inputs differ from the formula's: raw times, status flags and costs instead of censoring probabilities.
Inputs and outputs:
time: Follow-up time of each patient, truncated at the time limit; required, distinct values. Unit: years.;status: 1 for a complete case (death or reaching the limit), 0 for censored; required. Unit: none.;cost: Cumulative cost observed by each patient's time; required. Unit: currency.;S_c: Probability of remaining uncensored at each sorted time, before any censoring at that time. Unit: probability.;w: Weight of each sorted patient, status divided by S_c. Unit: none.;mu: Weighted mean cost. Unit: currency.;naive: Full-sample mean. Unit: currency.;complete: Mean over complete cases only. Unit: currency.Assumption: Censoring independent of costs and survival, follow-up truncated at a limit with patients alive at the limit counted as complete, and distinct times; with ties the order of deaths and censorings must be stated.
Worked example (Five patients over three years): Times of 0.5, 1, 1.5, 2 and 3 years, statuses 1, 0, 1, 0 and 1 and costs of 6,000, 5,000, 18,000, 9,000 and 12,000 pounds give censoring survival of 1, 1, 0.75, 0.75 and 0.375, weights of 1, 0, 1.333, 0 and 2.667 and a weighted mean of 12,400 pounds, against 10,000 for the full sample and 12,000 for complete cases only (computed here for illustration).
time = [0.5, 1, 1.5, 2, 3]; status = [1, 0, 1, 0, 1]; cost = [6000, 5000, 18000, 9000, 12000]; S_c = [1, 1, 0.75, 0.75, 0.375]; mu = 12400; naive = 10000; complete = 12000Worked example (No censoring): With every status 1 all weights are 1 and the weighted, full-sample and complete-case means all equal 10,000 pounds.
status = [1, 1, 1, 1, 1]; mu = 10000Excel: With patients sorted by time in rows 2 to 6, status in column B and the number at risk in column D (=COUNT(B$2:B$6)-ROW()+2), censoring survival in C2 is 1 and
=C2*(1-(1-B2)/D2)filled down from C3 gives each later value; the weighted mean is=SUMPRODUCT(B2:B6,E2:E6/C2:C6)/COUNT(E2:E6)with costs in column E.R:
bt_cost <- function(time, status, cost) { o <- order(time); st <- status[o]; A <- cost[o]; N <- length(A); Sc <- cumprod(c(1, 1-(1-st)/(N:1)))[1:N]; w <- st/Sc; list(S_c = Sc, w = w, mu = sum(w*A)/N, naive = mean(A), complete = mean(A[st == 1])) }Base R only;bt_cost(c(0.5, 1, 1.5, 2, 3), c(1, 0, 1, 0, 1), c(6000, 5000, 18000, 9000, 12000))returns mu = 12400.Python:
def bt_cost(time, status, cost): d = sorted(zip(time, status, cost)); N = len(d); Sc = list(itertools.accumulate([1-(1-st)/(N-i) for i, (t, st, c) in enumerate(d)], operator.mul, initial=1))[:N]; w = [st/s for (t, st, c), s in zip(d, Sc)]; return {"S_c": Sc, "w": w, "mu": sum(wi*c for wi, (t, st, c) in zip(w, d))/N, "naive": sum(c for t, st, c in d)/N, "complete": sum(c for t, st, c in d if st)/sum(st for t, st, c in d)}Needsimport itertools, operator; returns the same values as the R function.Test (Weights add up to the number of patients): With distinct times and the longest time a complete case, the weights add to N, 5 in the first example. Expected result: TRUE. Excel check:
=ABS(SUMPRODUCT(B2:B6,1/C2:C6)-COUNT(E2:E6))<1E-9Test (No censoring returns the plain mean): With every status 1 the weighted mean equals AVERAGE of the costs. Expected result: TRUE. Excel check, with WeightedMean holding the result:
=IF(MIN(B2:B6)=1,ABS(WeightedMean-AVERAGE(E2:E6))<1E-9,TRUE)Common error (Censoring probabilities estimated with deaths as the event): Using the ordinary survival curve in place of the censoring curve gives weights that do not add to N and a mean that has no justification; censoring must be the event.
Source: Wijeysundera HC, Wang X, Tomlinson G, Ko DT, Krahn MD. Techniques for estimating health care costs with censored data: an overview for the health services researcher. ClinicoEconomics and Outcomes Research. 2012;4:145-155. doi:10.2147/CEOR.S31552. Section Reweighted estimators, equation 2, and Kaplan-Meier estimates for survival and censoring; Bang H, Tsiatis AA. Estimating medical costs with censored data. Biometrika. 2000;87(2):329-343. doi:10.1093/biomet/87.2.329 (abstract read). Abstract.
S_c(t_(i)) = prod_(j < i) (1 - (1 - status_(j)) / (N - j + 1)); w_i = status_i / S_c(t_i); mu = sum_i (w_i * cost_i) / N
Try this function
Implementations
Excel
Weighted mean cost from named patient ranges
With 1 or 0 complete-case flags in Complete, costs in Costs and censoring survival in CensSurv (one row per patient, any positive value in censored rows), the formula returns the weighted mean, held in IPWMeanCost.
=SUMPRODUCT(Complete,Costs/CensSurv)/COUNT(Costs)
Assumptions
Censoring independent of the cost and survival process
Kaplan-Meier estimation of the censoring distribution assumes censoring does not depend on prognosis or spending.
Restricted period with enough complete cases for the weighted cost estimator
Costs are estimated up to a limit within follow-up, and S_c_i stays well above 0; with heavy censoring a few complete cases carry large weights.
Worked examples
Five patients with two censored over three years
Patients die at 0.5 and 1.5 years with costs of 6,000 and 18,000 pounds, are censored at 1 and 2 years with 5,000 and 9,000, and one is followed to the 3-year limit with 12,000. The censoring survival is 1, 0.75 and 0.375 at the three complete-case times, so the weighted mean is (6,000 + 24,000 + 32,000) / 5, or 12,400 pounds, against a full-sample mean of 10,000 (computed here for illustration).
N = 5; Delta_i = [1,0,1,0,1]; A_i = [6000,5000,18000,9000,12000]; S_c_i = [1,1,0.75,0.75,0.375]; mu = 12400
Same five patients with no censoring in the weighted cost estimator
With every patient a complete case, all weights are 1 and the estimator equals the plain mean of 10,000 pounds (computed here for illustration).
N = 5; Delta_i = [1,1,1,1,1]; A_i = [6000,5000,18000,9000,12000]; S_c_i = [1,1,1,1,1]; mu = 10000
Common errors
Averaging observed costs over all patients
The full-sample mean counts censored patients' partial costs as complete and is biased downwards; in heart failure data with 32 per cent censored over 1,080 days it gave 30,420 Canadian dollars against 36,490 from reweighted estimators.
Running a Kaplan-Meier analysis on cumulative cost
Treating cost as the time scale assumes cost at censoring is independent of cost at death, but with a constant spending rate R the two are R times the censoring time and R times the death time, so survival methods are not valid for cost.
Relying on the simple weighted estimator under heavy censoring
Only complete cases carry information, so the estimator becomes unstable as censoring grows; in simulations it substantially overestimated true costs with 53 per cent censoring.
Sources
Simple inverse probability weighted estimator of mean cost
Wijeysundera HC, Wang X, Tomlinson G, Ko DT, Krahn MD. Techniques for estimating health care costs with censored data: an overview for the health services researcher. ClinicoEconomics and Outcomes Research. 2012;4:145-155. doi:10.2147/CEOR.S31552. Section Reweighted estimators, equation 2: Bang and Tsiatis described an inverse probability weighted estimator that accommodates continuous censoring times; each uncensored participant is weighted by the inverse of the Kaplan-Meier estimate for censoring, S_c, defined as the probability of being uncensored beyond t with the roles of death and censoring reversed, and the mean total cost is 1/N times the sum of Delta_i A_i / S_c(t_i) over the full sample. Abstract: over a restricted 1,080 days with 32 per cent censored, the full-sample estimator gave 30,420 against 36,490 for the reweighted estimators, in 2008 Canadian dollars. Simulations: with heavy censoring (53 per cent) the simple estimator substantially overestimated true costs. Lifetime costs: cost to censoring and cost to death are both products of the same spending rate R and are not independent.
Weighted estimators of mean medical cost that account for censoring
Bang H, Tsiatis AA. Estimating medical costs with censored data. Biometrika. 2000;87(2):329-343. doi:10.1093/biomet/87.2.329 (abstract read). Abstract: naive analysis using summary statistics on censored cost data can give severely misleading inference; a class of weighted estimators that account appropriately for censoring is introduced for the mean medical cost of individuals whose costs may be right censored, shown to be consistent and asymptotically normal, and applied to cost data from a cardiology trial.
Canonical Identity
Stable URI · Machine-readable · Resolvable · CC BY 4.0