Ideal bootstrap standard error of a sample mean

For a sample mean the ideal bootstrap standard error, the value over every possible resample, has a closed form derived by Efron and Tibshirani. It equals the usual standard error, s divided by the square root of n, multiplied by the square root of n minus 1 over n, so the bootstrap understates the usual value in small samples.

Signature

x_bar = sum_(i=1)^n [x_i] / n; SE_inf = sqrt(sum_(i=1)^n [(x_i - x_bar)^2]) / n
Inputs
InputsDefinitionUnit
x_iValue for observation i, for example one patient's costunit of the data
nNumber of observations, for example patients in one armcount
Output
x_barMean of the n observationsunit of the data
SE_infIdeal bootstrap standard error of the sample mean, the limit of the Monte Carlo estimate as B growsunit of the data, for example pounds

Function

Bootstrap standard error of a trial statistic from resampled data

Maps a statistic and the data that produced it to the standard deviation of the statistic across samples drawn with replacement from the data, each of the same size as the original sample. The empirical distribution of the observations, with mass 1/n on each, stands in for the unknown population distribution, so no normal shape is assumed. In trial-based economic evaluation the statistic is usually a mean cost, an incremental cost or incremental net benefit, and the resampling copies the trial design by drawing patients within each arm with their costs and effects kept together.

Try this function

Implementations

  • Excel

    Ideal bootstrap standard error of a mean in one cell

    DEVSQ returns the sum of squared deviations from the mean for the range named Costs, and COUNT returns n.

    =SQRT(DEVSQ(Costs))/COUNT(Costs)

Assumptions

  • Independent observations from one arm

    The observations are independent draws from one arm or one population, and each bootstrap sample has the same size n.

  • Ideal value without simulation noise

    The formula gives the limit that the Monte Carlo estimate approaches as B grows, so it carries no simulation noise and serves as a check on a resampling routine.

Worked examples

  • Mean cost in arm T of the six-patient example

    Arm T costs have a mean of 2,000 pounds and squared deviations summing to 10,000,000, so the ideal bootstrap standard error is about 527.05 pounds, against 577.35 from s divided by the square root of n.

    x_i = [1000, 1200, 1400, 1600, 2000, 4800]; n = 6; x_bar = 2000; SE_inf = 527.05
  • Mean cost in arm U of the six-patient example

    Arm U costs have a mean of 1,000 pounds and squared deviations summing to 580,000, giving an ideal bootstrap standard error of about 126.93 pounds, against 139.04 from the usual formula.

    x_i = [600, 800, 900, 1000, 1100, 1600]; n = 6; x_bar = 1000; SE_inf = 126.93

Common errors

  • Reading the small-sample shortfall of the bootstrap as lower cost variability

    With six patients the bootstrap standard error is about 9% below the usual value, 527.05 against 577.35 pounds in arm T, before skewness plays any part. The shortfall comes from the divisor n, not from the costs being less variable; with 50 patients the factor is about 0.990.

Sources

  • Efron and Tibshirani closed form for the bootstrap standard error of a mean

    Efron B, Tibshirani R. Bootstrap methods for standard errors, confidence intervals, and other measures of statistical accuracy. Statistical Science. 1986;1(1):54-75. Section 1, equations 1.3 and 1.5 to 1.6, which give the usual estimate with divisor n minus 1 and the bootstrap estimate with divisor n, and note that the difference is too small to matter in most applications.

    View source →

Canonical Identity

Stable URI · Machine-readable · Resolvable · CC BY 4.0