Functions & Formulae

Each applied formula has its own function page, with a signature, implementations, and tests.

Schwarz criterion for ranking survival curves and approximating Bayes factors

BIC(lnL, k, c) = -2 * lnL + k * log(c); 2 * log(B_12) ≈ BIC_2 - BIC_1

Maps the maximised log-likelihood, the number of estimated parameters and a sample count for each model fitted to the same data to the Bayesian information criterion: minus twice the log-likelihood plus the parameter count times the natural log of the count. Lower values rank higher, and the difference between two models approximates twice the natural log of the Bayes factor between them, which also yields approximate posterior model probabilities. The patient-count form, with its worked examples, is recorded on the Akaike information criterion page; this page adds the event-count penalty for censored survival data and the Bayesian readings of BIC differences. Like AIC, the criterion describes fit within follow-up only.

  • BIC with the number of uncensored events as the penalty count

    BIC_d = -2 * lnL + k * log(d)

    Volinsky and Raftery's revision for censored survival data replaces the number of patients in the BIC penalty with the number of uncensored events d. With heavy censoring d is much smaller than the number of patients, so each parameter costs less and BIC moves closer to AIC. The function log is the natural logarithm.

  • Approximate Bayes factor from a difference in BIC

    B_12 = exp((BIC_2 - BIC_1) / 2)

    Kass and Raftery show that the Schwarz criterion is a rough approximation to the log Bayes factor, so twice the natural log of the Bayes factor in favour of model 1 over model 2 is approximately BIC_2 minus BIC_1. Exponentiating half the difference gives an approximate Bayes factor. On the Kass and Raftery scale, a difference of 0 to 2 is not worth more than a bare mention, 2 to 6 is positive, 6 to 10 strong and above 10 very strong evidence against the higher-BIC model.

  • Approximate posterior probability of a candidate model from BIC differences

    Delta_i = BIC_i - BIC_min; p_i = exp(-Delta_i / 2) / S

    Converts BIC differences into approximate posterior probabilities for each candidate model, assuming every model has the same prior probability. S is the sum of exp(-Delta_r / 2) over all R candidate models, so the probabilities add to 1 across the set. Burnham and Anderson view such model weights as a step towards model averaging.