Coefficient of variation of a bootstrap standard error from B replicates

Efron and Tibshirani's approximation splits the variability of a bootstrap standard error into the sampling variability of the ideal value and the simulation noise from using B replicates. The second term shrinks as B grows, which is why 50 to 200 replicates are adequate for a standard error in most situations.

Signature

CV_B = sqrt(CV_inf^2 + (E_delta + 2) / (4 * B))
Inputs
InputsDefinitionUnit
CV_infCoefficient of variation of the ideal bootstrap standard error, typically 0.10 to 0.30 in Efron and Tibshirani's accountproportion
E_deltaExpected excess kurtosis of the bootstrap distribution of the statistic; zero for a normal distributionnone
BNumber of bootstrap replicates used for the standard errorcount
Output
CV_BCoefficient of variation of the standard error computed from B replicatesproportion

Function

Bootstrap standard error of a trial statistic from resampled data

Maps a statistic and the data that produced it to the standard deviation of the statistic across samples drawn with replacement from the data, each of the same size as the original sample. The empirical distribution of the observations, with mass 1/n on each, stands in for the unknown population distribution, so no normal shape is assumed. In trial-based economic evaluation the statistic is usually a mean cost, an incremental cost or incremental net benefit, and the resampling copies the trial design by drawing patients within each arm with their costs and effects kept together.

Try this function

Implementations

  • Excel

    Coefficient of variation of a bootstrap standard error in one cell

    With named cells for the ideal coefficient of variation, the excess kurtosis and the number of replicates.

    =SQRT(CVIdeal^2+(Kurtosis+2)/(4*Replicates))

Assumptions

  • Large-sample approximation for the replicate count

    The expression is an approximation. Setting E_delta to zero corresponds to a normal bootstrap distribution; heavier tails raise the simulation term and the number of replicates needed.

Worked examples

  • One hundred replicates with an ideal coefficient of variation of 0.10

    With the kurtosis term at zero, 100 replicates raise the coefficient of variation from 0.10 to about 0.122, the article's figure, so the simulation adds little.

    CV_inf = 0.10; E_delta = 0; B = 100; CV_B = 0.1225
  • Fifty replicates at the lower end of the adequate range

    Fifty replicates give a coefficient of variation of about 0.141 when the ideal value is 0.10.

    CV_inf = 0.10; E_delta = 0; B = 50; CV_B = 0.1414
  • One hundred replicates with an ideal coefficient of variation of 0.30

    At the upper end of the typical range, 100 replicates move the coefficient of variation only from 0.30 to about 0.308.

    CV_inf = 0.30; E_delta = 0; B = 100; CV_B = 0.3082

Common errors

  • Setting the replicate count for an interval by the standard error rule

    The 50 to 200 replicates that suffice for a standard error are too few for bias-corrected intervals, for which Efron and Tibshirani give about 1,000 as a rough minimum, or for a percentile interval, about 250. The count is set by the most demanding output.

Sources

  • Efron and Tibshirani on the number of bootstrap replications

    Efron B, Tibshirani R. Bootstrap methods for standard errors, confidence intervals, and other measures of statistical accuracy. Statistical Science. 1986;1(1):54-75. Section 9, equation 9.1 and Table 10, which give the coefficient of variation approximation and the replicate numbers for intervals, and section 2, which judges 50 to 200 replicates adequate for a standard error.

    View source →

Canonical Identity