Approximate posterior probability of a candidate model from BIC differences

Converts BIC differences into approximate posterior probabilities for each candidate model, assuming every model has the same prior probability. S is the sum of exp(-Delta_r / 2) over all R candidate models, so the probabilities add to 1 across the set. Burnham and Anderson view such model weights as a step towards model averaging.

Signature

Delta_i = BIC_i - BIC_min; p_i = exp(-Delta_i / 2) / S
Inputs
InputsDefinitionUnit
BIC_iBIC of candidate model inone
BIC_minSmallest BIC among the R candidate models, all computed with the same penalty countnone
SSum of exp(-Delta_r / 2) over all R candidate models, including model inone
Output
Delta_iBIC of model i minus the smallest BIC in the set; 0 for the best-scoring modelBIC units
p_iApproximate posterior probability that model i is the quasi-true model in the candidate setprobability

Function

Schwarz criterion for ranking survival curves and approximating Bayes factors

Maps the maximised log-likelihood, the number of estimated parameters and a sample count for each model fitted to the same data to the Bayesian information criterion: minus twice the log-likelihood plus the parameter count times the natural log of the count. Lower values rank higher, and the difference between two models approximates twice the natural log of the Bayes factor between them, which also yields approximate posterior model probabilities. The patient-count form, with its worked examples, is recorded on the Akaike information criterion page; this page adds the event-count penalty for censored survival data and the Bayesian readings of BIC differences. Like AIC, the criterion describes fit within follow-up only.

Try this function

Implementations

  • Excel

    BIC posterior probability from a column of BIC values

    With the BIC of one model in BICValue and the BIC values of all candidates in BICRange, the formula returns that model's approximate posterior probability.

    =EXP(-(BICValue-MIN(BICRange))/2)/SUMPRODUCT(EXP(-(BICRange-MIN(BICRange))/2))

Assumptions

  • Equal prior probability for every candidate curve in the BIC probabilities

    The formula assumes that every candidate model has the same prior probability, 1 divided by R. Burnham and Anderson note that the probabilities rest on this assumption.

  • Probability of the quasi-true model within the set

    Each probability is best read as the probability that a model is the quasi-true model in the set, the simplest candidate closest to the truth, not that it is the true process. The probabilities depend on the candidate set and shift whenever a curve is added or dropped.

Worked examples

  • Posterior probability of the generalised gamma with 120 deaths

    The generalised gamma has the smallest event-count BIC, so its difference is 0 and its probability is 1 divided by the sum of 1.9205, about 0.521.

    BIC_i = 1056.962; BIC_min = 1056.962; Delta_i = 0; S = 1.9205; p_i = 0.521
  • Posterior probability of the Weibull with 120 deaths

    The Weibull is 0.613 behind the generalised gamma, giving a probability of about 0.383, so the data give similar support to the two leading curves.

    BIC_i = 1057.575; BIC_min = 1056.962; Delta_i = 0.613; S = 1.9205; p_i = 0.383
  • Posterior probability of the exponential with 120 deaths

    The exponential is 7.825 behind, leaving it a probability of about 0.0104.

    BIC_i = 1064.787; BIC_min = 1056.962; Delta_i = 7.825; S = 1.9205; p_i = 0.0104

Common errors

  • Reading a BIC posterior probability near 1 as a plausible extrapolation

    A probability near 1 means that one model stands out on the data within follow-up, not that it will extrapolate well. NICE DSU TSD 14 treats AIC and BIC as measures of internal validity only, so the extrapolated tail is judged on external data, clinical plausibility and expert judgement.

  • Comparing BIC posterior probabilities across different candidate sets

    Because the probabilities sum to 1 over the candidates, dropping the exponential from the 120-death example raises the generalised gamma from about 0.521 to about 0.526, and dropping the Weibull instead raises it to about 0.844. Probabilities from analyses with different candidates cannot be compared.

Sources

  • Burnham and Anderson on BIC posterior model probabilities

    Burnham KP, Anderson DR. Multimodel inference: understanding AIC and BIC in model selection. Sociological Methods & Research. 2004;33(2):261-304. Page 275, which gives the posterior model probabilities from BIC differences under equal prior probabilities of 1/R, and pages 278 to 279, which read them as probabilities of the quasi-true model in the set.

    View source →

  • NICE DSU TSD 14 on the internal validity of BIC-ranked survival curves

    Latimer N. NICE DSU Technical Support Document 14: Survival analysis for economic evaluations alongside clinical trials, extrapolation with patient-level data. Sheffield: Decision Support Unit, ScHARR, University of Sheffield; 2011, last updated March 2013. Section 3.5, which states that AIC and BIC address the internal validity of fitted models but not their external validity.

    View source →

Canonical Identity