Bayesian information criterion as a companion to AIC

Schwarz's criterion replaces the AIC penalty of 2 per parameter with the natural logarithm of the sample size, so the penalty grows with n. With 300 patients it is about 5.70 per parameter, and BIC leans towards simpler distributions. NICE DSU TSD 14 recommends presenting AIC and BIC, or other suitable tests of internal validity, when survival models are compared. The function log is the natural logarithm.

Signature

BIC = k * log(n) - 2 * lnL
Inputs
InputsDefinitionUnit
kNumber of parameters estimated when the model is fittedcount
nSample size used in the penalty, usually the number of patients in the data to which the model was fittedcount
lnLNatural logarithm of the maximised likelihood of the modelnone
Output
BICBayesian information criterion of the fitted model. Lower values are preferred, and only differences between models fitted to the same data are readnone

Function

Penalised-likelihood ranking of candidate survival models

Maps the maximised log-likelihood and the number of estimated parameters of each model fitted to the same data to an information criterion, a fit statistic with a penalty for each extra parameter. Lower values rank higher, and only differences between models fitted to the same observations carry meaning. In health technology assessment the candidates are usually parametric survival curves fitted to trial time-to-event data before extrapolation. The criteria measure fit within follow-up only, so they inform but do not settle the choice of curve.

Try this function

Implementations

  • Excel

    BIC in one cell for comparison with AIC

    Excel LN is the natural logarithm. With named cells LogLik, Params and SampleSize, the formula returns the BIC.

    =Params*LN(SampleSize)-2*LogLik

Assumptions

  • Choice of n for censored survival data

    The sample size in the penalty is a choice when data are censored. Volinsky and Raftery proposed the number of uncensored events rather than the number of patients, which lowers the penalty when censoring is heavy. The same n is used for every candidate.

  • Equal prior probabilities for posterior model probabilities

    BIC differences convert into approximate posterior model probabilities only when every candidate model has an equal prior probability.

Worked examples

  • BIC of the log-normal curve with 300 patients

    With n of 300, the two-parameter log-normal with a maximised log-likelihood of -603.2 has a BIC of about 1217.81, the lowest of the six curves.

    k = 2; n = 300; lnL = -603.2; BIC = 1217.81
  • BIC of the generalised gamma curve with 300 patients

    The three-parameter generalised gamma has a BIC of about 1222.71. Its distance behind the log-normal grows from 1.2 units on AIC to 4.90 units on BIC, while the two-parameter models keep their AIC gaps.

    k = 3; n = 300; lnL = -602.8; BIC = 1222.71

Common errors

  • Using Excel LOG instead of LN when computing BIC alongside AIC

    Excel LOG uses base 10 by default. With 300 patients it gives a penalty of about 2.48 per parameter instead of about 5.70, so BIC barely differs from AIC and loses its stronger preference for simpler distributions.

Sources

  • BIC formula and its relation to AIC

    Burnham KP, Anderson DR. Multimodel inference: understanding AIC and BIC in model selection. Sociological Methods & Research. 2004;33(2):261-304. Page 275, which gives BIC as minus twice the log-likelihood plus K times the log of n, and the posterior model probabilities under equal prior probabilities.

    View source →

  • Schwarz's derivation of BIC as the companion to AIC

    Schwarz G. Estimating the dimension of a model. Annals of Statistics. 1978;6(2):461-464. Abstract and main result, which derive the criterion as the Bayes solution for choosing among models of different dimension.

    View source →

  • BIC penalises extra parameters more than AIC in survival comparisons

    Latimer N. NICE DSU Technical Support Document 14: Survival analysis for economic evaluations alongside clinical trials, extrapolation with patient-level data. Sheffield: Decision Support Unit, ScHARR, University of Sheffield; 2011, last updated March 2013. Section 3.3, which notes that additional parameters are penalised more highly by BIC than by AIC, and section 5, which recommends presenting AIC and BIC tests.

    View source →

  • Number of events as the BIC sample size in survival comparisons with AIC

    Volinsky CT, Raftery AE. Bayesian information criterion for censored survival models. Biometrics. 2000;56(1):256-262. Abstract, which proposes defining the BIC penalty by the number of uncensored events rather than the number of observations.

    View source →

Canonical Identity

Stable URI · Machine-readable · Resolvable · CC BY 4.0