Signature
BS_max = o_bar * (1 - o_bar); BS_scaled = 1 - BS / BS_max
| Inputs | Definition | Unit |
|---|---|---|
o_bar | Overall observed event proportion in the validation sample | proportion |
BS | Brier score of the model in the same patients | none |
BS_max | Brier score of a model that gives every patient the overall event rate | none |
|---|---|---|
BS_scaled | Scaled Brier score, often reported as a percentage; negative when the model does worse than the non-informative model | proportion |
Function
Brier score as a proper scoring rule for predicted risks
Maps a set of predicted event probabilities and the observed binary outcomes to the mean squared difference between them, a proper scoring rule in which lower values are better. The score rewards predictions that are both well calibrated and able to separate patients who have the event from those who do not, and it can be partitioned into reliability, resolution and uncertainty terms. In health economic models it is one check on a risk equation before its predictions become event probabilities. Discrimination alone is covered by HE-FM-AUC-001.
Try this function
Implementations
Excel
Scaled Brier score in one cell
With the Brier score in a cell named Brier and the observed event rate in EventRate.
=1-Brier/(EventRate*(1-EventRate))
Assumptions
Reference based on the observed event rate
The reference uses the observed rate, so it is the same for every model compared in the same patients. When a model is correct on average, its mean predicted risk equals the observed rate and the two versions of the reference agree.
Event rate strictly inside the probability scale
The reference is zero, and the scaled score undefined, when no patient or every patient has the event.
Worked examples
Scaled score of model A in the admission example
Model A captures about 24% of the possible improvement over the non-informative model.
BS = 0.146; o_bar = 0.26; BS_max = 0.1924; BS_scaled = 0.2412
Scaled score of model B in the admission example
Against the same reference, the overpredicting model B captures about 15% of the possible improvement.
BS = 0.16375; o_bar = 0.26; BS_max = 0.1924; BS_scaled = 0.1489
Non-informative model scales to zero
The ten-patient non-informative model with a 10% event rate scores 0.09, equal to its reference, so its scaled score is 0.
BS = 0.09; o_bar = 0.10; BS_max = 0.09; BS_scaled = 0
Common errors
Scaling each model against its own mean predicted risk
Model B predicts a mean risk of 0.395 against an observed rate of 0.26. Scaled against its own mean prediction, its reference rises to about 0.239 and its scaled score to about 0.315, wrongly suggesting that it outperforms model A at 0.241. Against the observed rate it scores about 0.149.
Sources
Steyerberg and colleagues on the scaled Brier score
Steyerberg EW, Vickers AJ, Cook NR, Gerds T, Gonen M, Obuchowski N, Pencina MJ, Kattan MW. Assessing the performance of prediction models: a framework for traditional and novel measures. Epidemiology. 2010;21(1):128-138. Section on overall performance, which scales the score by its maximum under a non-informative model, written with the mean predicted risk, to range from 0% to 100%, and notes that it is very similar to Pearson's R2.
Canonical Identity
Stable URI · Machine-readable · Resolvable · CC BY 4.0