Akaike information criterion from the maximised log-likelihood

Akaike's criterion charges a penalty of 2 for each estimated parameter. A more complex model is therefore preferred only if each extra parameter raises the maximised log-likelihood by more than one unit. Because it compares maximised likelihoods rather than testing one model against another, it can rank survival distributions that are not nested within one another.

Signature

AIC = 2 * k - 2 * lnL
Inputs
InputsDefinitionUnit
kNumber of parameters estimated when the model is fitted, for example 1 for the exponential, 2 for the Weibull, Gompertz, log-logistic and log-normal, and 3 for the generalised gammacount
lnLNatural logarithm of the maximised likelihood of the model, as reported by the fitting software. For survival data it includes the contribution of censored observationsnone; usually a negative number
Output
AICAkaike information criterion of the fitted model. Lower values indicate a better balance of fit and parameter countnone; the value contains an arbitrary constant and is read only as a difference from other models fitted to the same data

Function

Penalised-likelihood ranking of candidate survival models

Maps the maximised log-likelihood and the number of estimated parameters of each model fitted to the same data to an information criterion, a fit statistic with a penalty for each extra parameter. Lower values rank higher, and only differences between models fitted to the same observations carry meaning. In health technology assessment the candidates are usually parametric survival curves fitted to trial time-to-event data before extrapolation. The criteria measure fit within follow-up only, so they inform but do not settle the choice of curve.

Computational function

  • Computational function: AIC, AIC differences and Akaike weights for candidate survival curves

    Takes the maximised log-likelihoods and parameter counts of a set of candidate survival models fitted to the same data, and returns three vectors: AIC, the AIC difference from the best-scoring model and the Akaike weight of each model. It chains HE-FM-AIC-001, HE-FM-AIC-002 and HE-FM-AIC-003, and the second and third steps use the whole vector produced by the step before, so the inputs are lists of models rather than the single values the formulae take.

    Inputs and outputs: lnL: Vector of maximised log-likelihoods, one per candidate model, as reported by the fitting software; required, all from the same patients, endpoint and data cut. Unit: none.; k: Vector of the number of estimated parameters of each model, in the same order; required, positive integers. Unit: count.; AIC: Vector of AIC values. Unit: none.; delta: Vector of AIC differences from the smallest AIC; 0 for the best-scoring model. Unit: AIC units.; w: Vector of Akaike weights, summing to 1. Unit: proportion from 0 to 1.

    Assumption: All candidates are fitted by maximum likelihood to identical data, and the set is specified before fitting. The weights are conditional on the set: adding or removing a model changes every weight. Subtracting the minimum before exponentiating also avoids exponentials of large negative numbers, which underflow to zero in floating-point arithmetic when AIC values run into the thousands.

    Worked example (Six candidate curves for one trial arm of 300 patients): The article's illustrative comparison, in the order exponential, Weibull, Gompertz, log-logistic, log-normal and generalised gamma. The log-normal ranks first with a weight of about 0.44, and the generalised gamma and log-logistic are within 2 units of it. lnL = (-612.4, -605.1, -606.0, -603.9, -603.2, -602.8); k = (1, 2, 2, 2, 2, 3); AIC = (1226.8, 1214.2, 1216.0, 1211.8, 1210.4, 1211.6); delta = (16.4, 3.8, 5.6, 1.4, 0, 1.2); w = (0.0001, 0.0663, 0.0270, 0.2201, 0.4433, 0.2433)

    Worked example (Extra parameter that buys exactly one log-likelihood unit): A two-parameter model whose log-likelihood is exactly 1 unit above a one-parameter model's ties with it on AIC, so each receives a weight of 0.5. This limiting case checks that the penalty and the weights are implemented correctly. lnL = (-100, -99); k = (1, 2); AIC = (202, 202); delta = (0, 0); w = (0.5, 0.5)

    Excel: =2*C2-2*B2 in D2, =D2-MIN($D$2:$D$7) in E2 and =EXP(-E2/2)/SUMPRODUCT(EXP(-$E$2:$E$7/2)) in F2, each filled down to row 7. With one model per row, log-likelihoods in B2:B7 and parameter counts in C2:C7, columns D, E and F return AIC, the AIC difference and the Akaike weight.

    R: akaike_table <- function(loglik, k) { aic <- 2*k-2*loglik; delta <- aic-min(aic); w <- exp(-delta/2)/sum(exp(-delta/2)); data.frame(aic, delta, w) } Vectorised over models; akaike_table(c(-612.4, -605.1, -606.0, -603.9, -603.2, -602.8), c(1, 2, 2, 2, 2, 3)) reproduces the six-curve example.

    Python: def akaike_table(loglik, k): aic = [2*ki-2*li for li, ki in zip(loglik, k)]; m = min(aic); delta = [a-m for a in aic]; e = [math.exp(-d/2) for d in delta]; s = sum(e); return aic, delta, [x/s for x in e] Uses the math module and returns the three lists in the order of the inputs.

    Test (Weights sum to one): The Akaike weights of all candidates add to 1. Expected result: TRUE. Excel check: =ABS(SUM(F2:F7)-1)<1E-9

    Test (Ratio of two weights does not depend on the other models): The ratio of the log-normal weight to the Weibull weight equals exp(1.9), about 6.69, whichever other models are in the set. Expected result: TRUE. Excel check: =ABS(F6/F3-EXP((E3-E6)/2))<1E-9

    Common error (Comparing weights from candidate sets of different size): Removing the exponential from the six-curve example leaves the log-normal weight at about 0.44 only because the exponential contributed almost nothing; removing the generalised gamma instead would raise it to about 0.59. Weights from different candidate sets cannot be compared, and the set should be reported with them.

    Source: Burnham KP, Anderson DR. Multimodel inference: understanding AIC and BIC in model selection. Sociological Methods & Research. 2004;33(2):261-304. Pages 271 to 272, which define the AIC differences, the likelihood of each model given the data and the Akaike weights, normalised to sum to 1 over the model set.

    AIC = 2 * k - 2 * lnL; delta = AIC - min(AIC); w = exp(-delta / 2) / sum(exp(-delta / 2))

Try this function

Implementations

  • Excel

    AIC in one cell from the log-likelihood and parameter count

    With the maximised log-likelihood in a cell named LogLik and the number of estimated parameters in a cell named Params, the formula returns the AIC.

    =2*Params-2*LogLik

Assumptions

  • Candidate models fitted to identical survival data

    Every candidate is fitted by maximum likelihood to the same patients, the same endpoint and the same data cut. A model fitted to a different subset or a later data cut has a likelihood for different observations, and its AIC cannot be compared with the others.

  • Parameter count includes every estimated parameter

    k counts all parameters estimated from the data. NICE DSU TSD 14 gives the exponential one parameter, the Weibull, Gompertz, log-logistic and log-normal two each, and the generalised gamma three. In spline-based models the count rises with the number of knots.

  • Candidate set chosen before fitting

    The set of candidate distributions is specified before the models are fitted, and it need not contain a true model. Burnham and Anderson treat a well-chosen set of models specified in advance as a precondition of the method.

Worked examples

  • AIC of the log-normal curve in a six-curve comparison

    In the article's illustrative example, a two-parameter log-normal curve fitted to one trial arm of 300 patients has a maximised log-likelihood of -603.2. Its AIC is 4 plus 1206.4, or 1210.4, the lowest of the six candidates.

    k = 2; lnL = -603.2; AIC = 1210.4
  • AIC of the generalised gamma curve in a six-curve comparison

    The three-parameter generalised gamma fits the same data better by 0.4 log-likelihood units, at -602.8, but pays 2 more in penalty. Its AIC of 1211.6 ranks it second, behind the log-normal.

    k = 3; lnL = -602.8; AIC = 1211.6
  • AIC of the exponential curve in a six-curve comparison

    The one-parameter exponential has the smallest penalty but the poorest fit, with a maximised log-likelihood of -612.4, giving an AIC of 1226.8.

    k = 1; lnL = -612.4; AIC = 1226.8

Common errors

  • Comparing AIC across models fitted to different data

    AIC values from curves fitted to different patients, endpoints or data cuts are not comparable, because each likelihood describes different observations. When treatment arms are fitted separately the rankings may disagree, and NICE DSU TSD 14 advises the same type of distribution in both arms unless there is strong evidence that an alternative is more plausible.

  • Reading the lowest AIC as evidence of plausible extrapolation

    AIC measures fit to the observed data only. In a simulated example in NICE DSU TSD 21, the log-normal had the lowest AIC, yet the two models with the lowest AIC gave the poorest 40-year extrapolations. The clinical plausibility of the extrapolated hazard has to be assessed separately.

Sources

  • Akaike's definition of the information criterion

    Akaike H. A new look at the statistical model identification. IEEE Transactions on Automatic Control. 1974;19(6):716-723. Abstract, which defines AIC as minus twice the maximised log-likelihood plus twice the number of independently adjusted parameters, with the minimum preferred.

    View source →

  • AIC as an estimate of relative Kullback-Leibler information

    Burnham KP, Anderson DR. Multimodel inference: understanding AIC and BIC in model selection. Sociological Methods & Research. 2004;33(2):261-304. Pages 268 to 271, on AIC as an estimate of relative expected Kullback-Leibler information, and page 275, on the need for a good set of candidate models chosen in advance.

    View source →

  • AIC for non-nested survival models in NICE appraisals

    Latimer N. NICE DSU Technical Support Document 14: Survival analysis for economic evaluations alongside clinical trials, extrapolation with patient-level data. Sheffield: Decision Support Unit, ScHARR, University of Sheffield; 2011, last updated March 2013. Section 2 (parameter counts of the six standard distributions), section 3.3 (AIC and BIC compare models that need not be nested) and section 3.5 (fit statistics address internal validity only).

    View source →

  • Limits of AIC for survival extrapolation

    Rutherford MJ, Lambert PC, Sweeting MJ, Pennington B, Crowther MJ, Abrams KR, Latimer NR. NICE DSU Technical Support Document 21: Flexible methods for survival analysis. Sheffield: Decision Support Unit, ScHARR, University of Sheffield; 2020 (updated March 2022). Sections 3.2.2 and 3.2.3 (AIC and BIC give little information about extrapolation) and section 6.3 (simulated example).

    View source →

Canonical Identity

Stable URI · Machine-readable · Resolvable · CC BY 4.0