Murphy partition of the Brier score into reliability, resolution and uncertainty

When predictions take a limited set of K distinct values, the Brier score splits exactly into a reliability term, which measures miscalibration, minus a resolution term, which measures how far the group event rates spread from the overall rate, plus an uncertainty term, the score of the non-informative model. Refinement, uncertainty minus resolution, is the average uncertainty left within the groups, and the score also equals reliability plus refinement.

Signature

REL = sum_(k=1)^K [n_k * (f_k - o_k)^2] / N; RES = sum_(k=1)^K [n_k * (o_k - o_bar)^2] / N; UNC = o_bar * (1 - o_bar); REF = UNC - RES; BS = REL - RES + UNC
Inputs
InputsDefinitionUnit
n_kNumber of patients whose predicted risk is f_kcount
f_kThe kth distinct predicted riskprobability
o_kObserved event proportion among the n_k patients given predicted risk f_kproportion
NTotal number of patients, the sum of the n_kcount
o_barOverall observed event proportion in the sampleproportion
Output
RELReliability term, the size-weighted mean squared gap between each predicted value and the observed rate in its group; zero when every group is perfectly calibratednone
RESResolution term, the size-weighted spread of the group event rates around the overall rate; larger is betternone
UNCUncertainty term, the score of a non-informative model, which depends only on the outcomesnone
REFRefinement, the average uncertainty left within the groups the model createsnone
BSBrier score, equal to the patient-level score when groups are formed from identical predicted valuesnone
  • K Number of distinct predicted values, or of bins when risks are grouped (count)

Function

Brier score as a proper scoring rule for predicted risks

Maps a set of predicted event probabilities and the observed binary outcomes to the mean squared difference between them, a proper scoring rule in which lower values are better. The score rewards predictions that are both well calibrated and able to separate patients who have the event from those who do not, and it can be partitioned into reliability, resolution and uncertainty terms. In health economic models it is one check on a risk equation before its predictions become event probabilities. Discrimination alone is covered by HE-FM-AUC-001.

Try this function

Implementations

  • Excel

    Brier partition terms from a grouped table

    With one row per distinct predicted value, ranges Count, Pred and ObsRate, and the overall event rate in EventRate, the four formulas return reliability, resolution, uncertainty and the rebuilt score.

    =SUMPRODUCT(Count,(Pred-ObsRate)^2)/SUM(Count); =SUMPRODUCT(Count,(ObsRate-EventRate)^2)/SUM(Count); =EventRate*(1-EventRate); =Reliability-Resolution+Uncertainty

Assumptions

  • Patients grouped by identical predicted values

    The partition is exact when patients are grouped by identical predicted values. With continuous predicted risks the predictions are grouped into bins first, and the reliability and resolution terms then depend on the number and width of the bins.

  • Uncertainty term shared by models validated in the same patients

    The uncertainty term depends only on the overall event rate, not on the model, so models validated in the same patients share it, and models that place patients in the same groups also share resolution and refinement.

Worked examples

  • Partition of model A in the 100-patient admission example

    The well calibrated model A has a reliability term of only 0.0007; resolution is 0.0471, uncertainty 0.1924 and refinement 0.1453, and the partition reproduces the directly computed score of 0.146.

    K = 3; n_k = [50, 30, 20]; f_k = [0.10, 0.30, 0.60]; o_k = [0.08, 0.30, 0.65]; o_bar = 0.26; N = 100; REL = 0.0007; RES = 0.0471; UNC = 0.1924; REF = 0.1453; BS = 0.146
  • Partition of model B with the same groups and higher risks

    Model B places the patients in the same groups with higher predicted risks, so resolution, uncertainty and refinement are unchanged and only reliability rises, to 0.01845, giving a score of 0.16375.

    K = 3; n_k = [50, 30, 20]; f_k = [0.20, 0.45, 0.80]; o_k = [0.08, 0.30, 0.65]; o_bar = 0.26; N = 100; REL = 0.01845; RES = 0.0471; UNC = 0.1924; REF = 0.1453; BS = 0.16375

Common errors

  • Taking a shared AUC as evidence of equal calibration

    Models A and B rank the patients identically, so they share an AUC of about 0.788, yet model B's reliability term is 0.01845 against 0.0007 because it overpredicts admissions in every group. Only the reliability term exposes the difference.

  • Reporting partition terms without the binning used

    With continuous predictions the partition is computed after grouping into bins, and the reliability and resolution terms change with the number and width of the bins. Terms reported without the binning cannot be reproduced or compared.

Sources

  • Murphy vector partition of the probability score

    Murphy AH. A new vector partition of the probability score. Journal of Applied Meteorology. 1973;12(4):595-600. Abstract, which describes the three terms: the uncertainty inherent in the events, the reliability of the forecasts and the resolution of the forecasts.

    View source →

  • Bröcker decomposition of proper scores into reliability and resolution

    Bröcker J. Reliability, sufficiency, and the decomposition of proper scores. Quarterly Journal of the Royal Meteorological Society. 2009;135(643):1512-1519. Section 2, equations 10 and 13 with Table 1, which give the general partition and, for the Brier score, the squared difference as the divergence and p(1 minus p) as the entropy.

    View source →

  • Blattenberger and Lad calibration and refinement components

    Blattenberger G, Lad F. Separating the Brier score into calibration and refinement components: a graphical exposition. The American Statistician. 1985;39(1):26-32. The two-term form of the score as calibration plus refinement.

    View source →

Canonical Identity