Scaled Brier score against the non-informative model

Scales the Brier score against a non-informative model that predicts the overall event rate for every patient, so that 0 means no improvement on that model and 1 means perfect prediction. Steyerberg and colleagues write the reference in terms of the mean predicted risk; basing it on the observed rate keeps it fixed when models that are miscalibrated in the large are compared. A negative value means the model does worse than predicting the overall rate.

Signature

BS_max = o_bar * (1 - o_bar); BS_scaled = 1 - BS / BS_max
Inputs
InputsDefinitionUnit
o_barOverall observed event proportion in the validation sampleproportion
BSBrier score of the model in the same patientsnone
Output
BS_maxBrier score of a model that gives every patient the overall event ratenone
BS_scaledScaled Brier score, often reported as a percentage; negative when the model does worse than the non-informative modelproportion

Function

Brier score as a proper scoring rule for predicted risks

Maps a set of predicted event probabilities and the observed binary outcomes to the mean squared difference between them, a proper scoring rule in which lower values are better. The score rewards predictions that are both well calibrated and able to separate patients who have the event from those who do not, and it can be partitioned into reliability, resolution and uncertainty terms. In health economic models it is one check on a risk equation before its predictions become event probabilities. Discrimination alone is covered by HE-FM-AUC-001.

Try this function

Implementations

  • Excel

    Scaled Brier score in one cell

    With the Brier score in a cell named Brier and the observed event rate in EventRate.

    =1-Brier/(EventRate*(1-EventRate))

Assumptions

  • Reference based on the observed event rate

    The reference uses the observed rate, so it is the same for every model compared in the same patients. When a model is correct on average, its mean predicted risk equals the observed rate and the two versions of the reference agree.

  • Event rate strictly inside the probability scale

    The reference is zero, and the scaled score undefined, when no patient or every patient has the event.

Worked examples

  • Scaled score of model A in the admission example

    Model A captures about 24% of the possible improvement over the non-informative model.

    BS = 0.146; o_bar = 0.26; BS_max = 0.1924; BS_scaled = 0.2412
  • Scaled score of model B in the admission example

    Against the same reference, the overpredicting model B captures about 15% of the possible improvement.

    BS = 0.16375; o_bar = 0.26; BS_max = 0.1924; BS_scaled = 0.1489
  • Non-informative model scales to zero

    The ten-patient non-informative model with a 10% event rate scores 0.09, equal to its reference, so its scaled score is 0.

    BS = 0.09; o_bar = 0.10; BS_max = 0.09; BS_scaled = 0

Common errors

  • Scaling each model against its own mean predicted risk

    Model B predicts a mean risk of 0.395 against an observed rate of 0.26. Scaled against its own mean prediction, its reference rises to about 0.239 and its scaled score to about 0.315, wrongly suggesting that it outperforms model A at 0.241. Against the observed rate it scores about 0.149.

Sources

  • Steyerberg and colleagues on the scaled Brier score

    Steyerberg EW, Vickers AJ, Cook NR, Gerds T, Gonen M, Obuchowski N, Pencina MJ, Kattan MW. Assessing the performance of prediction models: a framework for traditional and novel measures. Epidemiology. 2010;21(1):128-138. Section on overall performance, which scales the score by its maximum under a non-informative model, written with the mean predicted risk, to range from 0% to 100%, and notes that it is very similar to Pearson's R2.

    View source →

Canonical Identity

Stable URI · Machine-readable · Resolvable · CC BY 4.0

Scaled Brier score against the non-informative model | HealthEconomics.wiki