Signature
BS = sum_(i=1)^N [(p_i - y_i)^2] / N
| Inputs | Definition | Unit |
|---|---|---|
p_i | Predicted probability of the event for patient i | probability |
y_i | Observed outcome for patient i: 1 if the event occurred and 0 otherwise | none |
N | Number of patients with a prediction and a known outcome | count |
BS | Brier score of the predictions; lower values are better | none |
|---|
Function
Brier score as a proper scoring rule for predicted risks
Maps a set of predicted event probabilities and the observed binary outcomes to the mean squared difference between them, a proper scoring rule in which lower values are better. The score rewards predictions that are both well calibrated and able to separate patients who have the event from those who do not, and it can be partitioned into reliability, resolution and uncertainty terms. In health economic models it is one check on a risk equation before its predictions become event probabilities. Discrimination alone is covered by HE-FM-AUC-001.
Computational function
Computational function: Brier score, Murphy partition and scaled score from predicted risks and outcomes
Takes a vector of predicted risks and a vector of observed binary outcomes, one entry per patient, and returns the Brier score, its Murphy partition and the scaled score. It applies HE-FM-BRIER-001 to the vectors, groups patients by identical predicted value to build the counts and observed rates that HE-FM-BRIER-002 needs, and then scales the score with HE-FM-BRIER-003, so the inputs are patient-level vectors rather than the group summaries the formulae take.
Inputs and outputs:
p: Vector of predicted event probabilities, one per patient; required, each a probability, not a percentage. Unit: probability.;y: Vector of observed outcomes in the same patient order, 1 for an event and 0 otherwise; required. Unit: none.;BS: Brier score. Unit: none.;REL: Reliability term of the partition. Unit: none.;RES: Resolution term of the partition. Unit: none.;UNC: Uncertainty term, the observed event rate times one minus it; BS equals REL minus RES plus UNC. Unit: none.;BS_scaled: Scaled Brier score against the observed event rate. Unit: proportion.Assumption: Grouping by identical predicted values makes the partition exact. With continuous risks every patient forms a group of one, the resolution term then equals the uncertainty term and the partition says nothing useful, so continuous risks are binned before the partition is read, and the bins are reported.
Worked example (Model A in the 100-patient admission example): The article's illustrative validation sample, written in R notation: 50, 30 and 20 patients at predicted risks of 0.10, 0.30 and 0.60, with 4, 9 and 13 admissions.
p = rep(c(0.10, 0.30, 0.60), c(50, 30, 20)); y = rep(c(1, 0, 1, 0, 1, 0), c(4, 46, 9, 21, 13, 7)); BS = 0.146; REL = 0.0007; RES = 0.0471; UNC = 0.1924; BS_scaled = 0.2412Worked example (Model B with the same patients): Model B ranks the patients in the same order with higher risks, so only the reliability term and therefore the score change.
p = rep(c(0.20, 0.45, 0.80), c(50, 30, 20)); y = rep(c(1, 0, 1, 0, 1, 0), c(4, 46, 9, 21, 13, 7)); BS = 0.16375; REL = 0.01845; RES = 0.0471; UNC = 0.1924; BS_scaled = 0.1489Excel:
=SUMXMY2(Predicted,Outcome)/COUNT(Predicted)returns the score from patient-level columns named Predicted and Outcome, and=1-(SUMXMY2(Predicted,Outcome)/COUNT(Predicted))/(AVERAGE(Outcome)*(1-AVERAGE(Outcome)))returns the scaled score. The partition terms come from a grouped table, as in HE-IM-BRIER-002.R:
brier_partition <- function(p, y) { n <- length(y); obar <- mean(y); nk <- tapply(y, p, length); ok <- tapply(y, p, mean); fk <- as.numeric(names(nk)); bs <- mean((p-y)^2); rel <- sum(nk*(fk-ok)^2)/n; res <- sum(nk*(ok-obar)^2)/n; unc <- obar*(1-obar); c(bs = bs, rel = rel, res = res, unc = unc, scaled = 1-bs/unc) }tapply groups the outcomes by distinct predicted value; the example vectors above can be pasted into R as written.Python:
def brier_partition(p, y): n = len(y); obar = sum(y)/n; g = {f: [yi for pi, yi in zip(p, y) if pi == f] for f in set(p)}; bs = sum((pi-yi)**2 for pi, yi in zip(p, y))/n; rel = sum(len(v)*(f-sum(v)/len(v))**2 for f, v in g.items())/n; res = sum(len(v)*(sum(v)/len(v)-obar)**2 for v in g.values())/n; unc = obar*(1-obar); return bs, rel, res, unc, 1-bs/uncPlain Python with no imports; the dictionary g holds the outcomes for each distinct predicted value.Test (Partition rebuilds the score): Reliability minus resolution plus uncertainty equals the directly computed Brier score. Expected result: TRUE. Excel check:
=ABS((Reliability-Resolution+Uncertainty)-Brier)<1E-9Test (Recalibrating to the group rates removes the reliability term): Replacing each patient's predicted risk with the observed rate in that patient's group, in a column named GroupRate, gives a score equal to the refinement, uncertainty minus resolution, which is 0.1453 in the example. Expected result: TRUE. Excel check:
=ABS(SUMXMY2(GroupRate,Outcome)/COUNT(Outcome)-(Uncertainty-Resolution))<1E-9Common error (Passing percentages instead of probabilities): Entering model A's risks as 10, 30 and 60 instead of 0.10, 0.30 and 0.60 gives a score of about 1,018 instead of 0.146. The predictions must be on the probability scale before the function is applied.
Source: Steyerberg EW, Vickers AJ, Cook NR, Gerds T, Gonen M, Obuchowski N, Pencina MJ, Kattan MW. Assessing the performance of prediction models: a framework for traditional and novel measures. Epidemiology. 2010;21(1):128-138. Section on overall performance, which defines the score and its scaled form. Murphy AH. A new vector partition of the probability score. Journal of Applied Meteorology. 1973;12(4):595-600. Abstract, on the three-term partition.
BS = mean((p - y)^2); REL = sum_k(n_k * (f_k - o_k)^2) / N; RES = sum_k(n_k * (o_k - o_bar)^2) / N; UNC = o_bar * (1 - o_bar); BS_scaled = 1 - BS / UNC
Try this function
Implementations
Excel
Brier score from prediction and outcome columns
SUMXMY2 returns the sum of squared differences between corresponding values in the two ranges, and dividing by the count gives the mean.
=SUMXMY2(Predicted,Outcome)/COUNT(Predicted)
Assumptions
Binary outcome known for every patient at the horizon
Each patient has a known binary outcome over the prediction horizon. For time-to-event outcomes with censoring, the score is computed at fixed time points with weighting for censoring and compared with a Kaplan-Meier benchmark, which this formula does not include.
Brier score as a property of a model in one population
The score depends on the prevalence of the outcome in the population where it is measured. Apparent performance in the development data is optimistic, so the score is estimated in an external sample or corrected with cross-validation or bootstrapping.
Worked examples
Non-informative prediction at a 10% event rate
Ten patients, one of whom has the event, all receive the overall event rate of 0.10. The score of 0.09 is the article's reference value for a non-informative model at a 10% event rate.
p_i = [0.1, 0.1, 0.1, 0.1, 0.1, 0.1, 0.1, 0.1, 0.1, 0.1]; y_i = [1, 0, 0, 0, 0, 0, 0, 0, 0, 0]; N = 10; BS = 0.09
Non-informative prediction at a 50% event rate
Two patients, one with the event, both given a risk of 0.5, score 0.25, the non-informative value at a 50% event rate.
p_i = [0.5, 0.5]; y_i = [1, 0]; N = 2; BS = 0.25
Confidently and completely wrong risk predictions
A risk of 1 for the patient without the event and 0 for the patient with it gives the worst possible score of 1.
p_i = [1, 0]; y_i = [0, 1]; N = 2; BS = 1
Common errors
Comparing raw Brier scores across populations with different event rates
A non-informative model scores 0.25 at a 50% event rate but 0.090 at a 10% rate, so a raw score from a low-risk primary care population can look better than one from a high-risk hospital population without any difference in model quality. The scaled score on this page corrects for the event rate.
Reading a low Brier score as evidence of clinical value
A lower score indicates more accurate probabilities but does not show whether using the model leads to better decisions. Assel, Sjoberg and Vickers found that when prevalence was low the score could favour a highly specific test where high sensitivity was needed, and recommended decision-analytic measures such as net benefit.
Sources
Steyerberg framework definition of the Brier score
Steyerberg EW, Vickers AJ, Cook NR, Gerds T, Gonen M, Obuchowski N, Pencina MJ, Kattan MW. Assessing the performance of prediction models: a framework for traditional and novel measures. Epidemiology. 2010;21(1):128-138. Section on overall performance, which defines the score as the mean squared difference between outcome and prediction and gives the non-informative values of 0.25 at 50% incidence and 0.090 at 10%.
Assel and colleagues on the Brier score as mean squared prediction error
Assel M, Sjoberg DD, Vickers AJ. The Brier score does not evaluate the clinical utility of diagnostic tests or prediction models. Diagnostic and Prognostic Research. 2017;1:19. Background and the section on the Brier score, which define it as the mean squared prediction error, state that it is a proper scoring rule and read its square root as the expected distance between prediction and outcome.
Brier's original verification score for probability forecasts
Brier GW. Verification of forecasts expressed in terms of probability. Monthly Weather Review. 1950;78(1):1-3. The paper that introduced the score to verify weather forecasts expressed as probabilities.
Microsoft documentation of the SUMXMY2 function
Microsoft. SUMXMY2 function. Microsoft Support. Description, which states that the function returns the sum of squares of differences of corresponding values in two arrays.
Canonical Identity
Stable URI · Machine-readable · Resolvable · CC BY 4.0