VerifiedEvidence: highv1.0.0

Model Selection

The process of choosing among alternative model structures or specifications for an evaluation, guided by fit, clinical plausibility, and parsimony.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

Concept

Theoretically, Model Selection is the process of identifying the statistical or mathematical model that provides the most appropriate representation of observed data for a specified analytical objective. It balances model fit, predictive performance, complexity and interpretability to avoid both underfitting and overfitting. The concept is grounded in statistical decision theory and information theory, recognising that several competing models may explain the same data but differ in their ability to generalise beyond the sample. In health economics, model selection is fundamental when choosing regression models, survival models, risk equations and other statistical models that provide inputs to economic evaluations.

Mathematically, model selection compares competing candidate models using recognised quantitative criteria. Depending on the modelling framework, comparisons may be based on likelihood functions, information criteria, goodness-of-fit measures, cross-validation error or formal hypothesis tests. Information criteria, particularly Akaike Information Criterion (AIC) and Bayesian Information Criterion (BIC), are widely used because they explicitly balance model fit against model complexity through a penalty for additional parameters.

In practice, model selection begins by specifying plausible candidate models based on subject-matter knowledge and theoretical considerations. Each model is estimated and evaluated using appropriate statistical criteria together with diagnostic assessment and validation procedures. Health economists typically combine quantitative measures with clinical plausibility, external evidence and intended model application before selecting the final statistical model for estimating costs, utilities, transition probabilities or treatment effects.


Purpose

Used to identify the statistical model that provides the most appropriate balance between explanatory performance, predictive accuracy and model complexity for health economic analysis.


Mathematical Formulae

Primary Formula

There is no universally recognised canonical mathematical formula.

Supporting Formulae

Akaike Information Criterion:

AIC = ?2ln(L) + 2k

Bayesian Information Criterion:

BIC = ?2ln(L) + k ln(n)

Likelihood Ratio Test:

LR = ?2[ln(L?) ? ln(L?)]

where:

  • L = maximum likelihood
  • k = number of estimated parameters
  • n = sample size
  • L? = likelihood of the restricted model
  • L? = likelihood of the unrestricted model

Related Mathematical Methods

  • Akaike Information Criterion
  • Bayesian Information Criterion
  • Likelihood Ratio Test
  • Cross-validation
  • Goodness of fit assessment
  • Residual analysis
  • Maximum likelihood estimation
  • Information-theoretic model comparison

Example

A health economist compares three regression models predicting annual healthcare costs.

ModelAICBIC
Linear regression2,4812,510
Generalised linear model2,4262,461
Generalised additive model2,4192,488

Although the generalised additive model has the lowest AIC, the generalised linear model has a substantially lower BIC than the more complex model and demonstrates satisfactory diagnostic performance. After considering predictive accuracy, model complexity and interpretability, the generalised linear model is selected for the economic evaluation.


Excel Implementation

FunctionExample FormulaHealth Economics Application
MIN=MIN(B2:B4)Identifies the model with the lowest information criterion.
MATCH=MATCH(MIN(B2:B4),B2:B4,0)Locates the best-performing model.
INDEX=INDEX(A2:A4,MATCH(MIN(B2:B4),B2:B4,0))Returns the selected model name.
RANK=RANK(B2,$B$2:$B$4,1)Ranks competing models using AIC or BIC values.

VBA (Optional)

Automate comparison of multiple candidate models by calculating selection criteria, ranking models and generating a model selection summary report.


Sources

  • Drummond MF, Sculpher MJ, Claxton K, Stoddart GL, Torrance GW. Methods for the Economic Evaluation of Health Care Programmes. 4th ed.
  • Briggs A, Claxton K, Sculpher M. Decision Modelling for Health Economic Evaluation.
  • Burnham KP, Anderson DR. Model Selection and Multimodel Inference. 2nd ed.
  • Claeskens G, Hjort NL. Model Selection and Model Averaging.
  • Harrell FE. Regression Modeling Strategies. 2nd ed.
  • Akaike H. A New Look at the Statistical Model Identification. IEEE Transactions on Automatic Control. 1974;19:716?723.

Library

Publications

1
  • Journal article

    Modeling Good Research Practices — Overview: A Report of the ISPOR-SMDM Modeling Good Research Practices Task Force-1 — Caro, Briggs, Siebert & Kuntz, Task Force Report 1 ed., 2012 (Value in Health / Medical Decision Making)

    The overview paper of the seven-part ISPOR-SMDM modelling good-practice series, setting out best-practice recommendations across model design, technique selection, implementation, validation, parameterisation, uncertainty and use in decision making.

Frequently Asked Questions (6)

  • What is model selection?

    The process of choosing among alternative model structures or specifications for an evaluation, guided by fit, clinical plausibility, and parsimony.

    Source: Philips et al. 2004

  • How are competing considerations balanced in model selection?

    Choosing among candidate models weighs several considerations that can pull in different directions. Closeness of fit to the data favours more complex models, parsimony favours simpler ones, and clinical plausibility requires that the chosen structure make sense for the disease whatever the statistics say. Selection balances these, preferring a model that fits adequately, remains as simple as the question allows, and represents the condition credibly. No single criterion decides it alone. Jackson and colleagues (2011) discuss this balance.

    Source: Jackson et al. 2011

  • What guides model selection?

    Model selection is guided by how well each candidate fits the relevant data, whether its structure and assumptions are clinically plausible and appropriate for the decision, and by parsimony, favouring the simplest specification that captures what matters. Fit indicates agreement with data, plausibility that the model makes clinical sense, and parsimony that it is no more complex than necessary. These considerations are weighed together, since a model should fit adequately, represent the problem credibly, and avoid unnecessary complexity, so no single criterion decides the choice alone.

    Source: Akaike 1974

  • How does fit inform model selection?

    Fit informs model selection by showing how well each candidate matches the observed data, favouring specifications that agree with it, and statistical criteria that balance fit against complexity, such as the Akaike or Bayesian information criteria, help compare candidates. However, fit is not the sole basis, since better fit can reflect overfitting, and a well-fitting model may still be implausible. Fit therefore contributes to selection alongside clinical plausibility and parsimony, informing but not determining the choice of structure.

    Source: Akaike 1974

  • Why does the choice of model structure matter?

    The choice of model structure matters because different structures can represent the problem differently and produce different results, so the selection can materially affect the conclusion. A structure that omits important features or misrepresents the disease will give biased results, while an unnecessarily complex one adds burden without benefit. Because structural choices can influence results as much as parameters, and are harder to correct later, selecting an appropriate structure, and justifying it, is fundamental to a valid evaluation, which is why model selection is made carefully.

    Source: Philips et al. 2004

  • How does parsimony feature in model selection?

    Parsimony features in model selection as a preference for the simplest structure that adequately captures what matters for the decision, so that among candidates fitting and representing the problem acceptably, the less complex is favoured. Parsimony guards against unnecessary complexity, which adds data demands, opacity, and error without improving the answer. In selection, parsimony is balanced against fit and plausibility, so the chosen model is realistic enough to answer the question validly while avoiding excess detail, embodying the principle of being as simple as possible but no simpler.

    Source: Roberts et al. 2012

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 16 Oct 2025

Content version: 1.0.0

Canonical Identity

Term code
HE-EM-MV-051

Stable URI · Machine-readable · Resolvable · CC BY 4.0