VerifiedEvidence: highv1.0.0

Calibration Assessment

An evaluation of how closely a calibrated model's predicted outputs match the observed data used as calibration targets, typically using formal fit statistics.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

Concept


Theoretically, Calibration Assessment is the evaluation of the agreement between predicted outcomes generated by a statistical or health economic model and the outcomes actually observed. It is founded on statistical prediction theory and model validation. The concept exists because an accurate prediction model should produce probabilities that correspond closely to observed event frequencies across the full range of predicted risk.

Mathematically, calibration is assessed by comparing predicted probabilities with observed event rates using calibration statistics, goodness-of-fit measures and graphical methods. The recognised mathematical framework includes calibration intercepts, calibration slopes, observed-to-expected ratios and goodness-of-fit tests such as the Hosmer?Lemeshow test. Perfect calibration occurs when predicted probabilities equal observed probabilities.

In practice, calibration assessment is performed during model validation by comparing model predictions with external or internal validation datasets. Calibration plots, observed-to-expected ratios and statistical tests are used to determine whether recalibration is required before a model is applied to health economic evaluation, risk prediction or decision modelling.

Purpose


Used to evaluate whether predicted risks or outcomes generated by a model accurately reflect observed outcomes, thereby assessing the reliability of prediction models used in health economic evaluation and decision making.

Mathematical Formulae

Primary Formula

O:E = O / E

where:

O = observed number of events

E = expected number of events

Perfect calibration:

O:E = 1

Supporting Formulae

Calibration model:

logit(Y) = � + ? ? logit(P)

Perfect calibration:

� = 0

? = 1

Hosmer?Lemeshow statistic:

?� = ? (O? ? E?)� / E?

Related Mathematical Methods

  • Calibration Plot
  • Hosmer?Lemeshow Test
  • Logistic Regression
  • Observed-to-Expected Ratio
  • Internal Validation
  • External Validation
  • Bootstrapping

Example

A cardiovascular risk model predicts 100 myocardial infarctions within a validation cohort over five years. Follow-up identifies 95 observed events.

Observed-to-expected ratio:

O:E = 95 / 100 = 0.95

An O:E ratio close to one indicates good calibration. The calibration intercept is estimated at ?0.04 and the calibration slope at 0.98, suggesting minimal systematic overprediction and excellent agreement between predicted and observed outcomes.


Excel Implementation

FunctionExample FormulaHealth Economics Application
SUM=SUM(B2:B101)Calculate total expected events
COUNTIFS=COUNTIFS(C2:C101,1)Calculate observed events
IF=IF(B105/C105>0.95,"Well Calibrated","Review Model")Assess observed-to-expected agreement
RSQ=RSQ(B2:B101,C2:C101)Assess agreement between predicted and observed outcomes

VBA (Optional)

Automate calibration plots, observed-to-expected calculations and validation summaries across multiple prediction models.


Sources

  • Harrell FE. Regression Modeling Strategies.
  • Steyerberg EW. Clinical Prediction Models.
  • Hosmer DW, Lemeshow S, Sturdivant RX. Applied Logistic Regression.
  • Riley RD, et al. Minimum requirements for the development and validation of clinical prediction models.
  • NICE. Health Technology Evaluations: The Manual.
  • ISPOR-SMDM Modeling Good Research Practices Task Force Reports.

Library

Publications

1
  • Journal article

    Modeling Good Research Practices — Overview: A Report of the ISPOR-SMDM Modeling Good Research Practices Task Force-1 — Caro, Briggs, Siebert & Kuntz, Task Force Report 1 ed., 2012 (Value in Health / Medical Decision Making)

    The overview paper of the seven-part ISPOR-SMDM modelling good-practice series, setting out best-practice recommendations across model design, technique selection, implementation, validation, parameterisation, uncertainty and use in decision making.

Frequently Asked Questions (6)

  • What is calibration assessment?

    An evaluation of how closely a calibrated model's predicted outputs match the observed data used as calibration targets, typically using formal fit statistics.

    Source: Eddy DM, Hollingworth W, Caro JJ, et al. Model transparency and validation: a report of the ISPOR-SMDM Modeling Good Research Practices Task Force-7. Value in Health. 2012;15(6):843-850. doi:10.1016/j.jval.2012.04.012.

  • How is the fit of a calibrated model judged?

    Once a model has been calibrated, its predicted outputs are compared against the real-world targets it was tuned to reproduce, and the closeness of that agreement is measured. Formal fit statistics quantify the gap between prediction and target, and a good calibration shows small, unbiased differences across the targets. Judging fit this way reveals whether the calibration succeeded or whether the model cannot match the data even after adjustment, which would signal a structural problem. Vanni and colleagues (2011) describe assessing fit.

    Source: Vanni et al. 2011

  • How is calibration assessment carried out?

    Calibration assessment is carried out by comparing the calibrated model's predicted outputs with the observed calibration targets, using goodness-of-fit measures or graphical comparison to quantify how closely they agree. The assessment examines whether the model reproduces the targets within acceptable bounds and where discrepancies remain. It may consider the uncertainty in the targets and in the calibrated parameters. The result indicates how well the calibration achieved agreement with the data, showing whether the model is adequately consistent with the outcomes it was calibrated to.

    Source: Eddy et al. 2012

  • Why is calibration assessment important?

    Calibration assessment is important because calibration is only useful if it actually makes the model reproduce the target data, and assessing the fit confirms whether it did. A poor fit indicates that the model, even after calibration, does not match the observed outcomes, signalling structural problems or unattainable targets. Good fit supports confidence that the calibrated parameters make the model consistent with reality. Assessing the calibration thus verifies the fitting step and reveals limitations, informing whether the calibrated model can be relied upon.

    Source: Eddy et al. 2012

  • What does poor calibration fit indicate?

    A poor calibration fit indicates that the model, even with adjusted parameters, cannot reproduce the observed target data well, which may point to problems with the model's structure, inappropriate assumptions, targets that are inconsistent or unattainable, or an inadequate calibration search. It signals that the model does not match reality in the respects measured, so its predictions may be unreliable. Identifying poor fit prompts investigation of the causes and possible revision of the model, rather than accepting a calibration that fails to match the data.

    Source: Eddy et al. 2012

  • How does calibration assessment relate to validation?

    Calibration assessment relates to validation but concerns the fit to the calibration targets specifically, whereas validation more broadly checks the model against data, ideally independent of the calibration. A good calibration fit shows the model reproduces the data it was fitted to, which is necessary but not sufficient for validity, since matching the calibration targets does not guarantee accuracy on other outcomes. Genuine validation compares the model with data not used in calibration, so calibration assessment is one part of establishing a model's credibility.

    Source: Briggs, Claxton & Sculpher 2006

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 13 Oct 2025

Content version: 1.0.0

Canonical Identity

Term code
HE-EM-MV-007

Stable URI · Machine-readable · Resolvable · CC BY 4.0