Signature
AUC = (n_conc + 0.5 * n_tie) / (n_1 * n_0)
| Inputs | Definition | Unit |
|---|---|---|
n_conc | Number of pairs, each of one person with the outcome and one without, in which the person with the outcome has the higher score | count of pairs |
n_tie | Number of pairs of one person with the outcome and one without in which the two scores are equal | count of pairs |
n_1 | Number of people with the outcome (cases) | count of people |
n_0 | Number of people without the outcome (non-cases) | count of people |
AUC | Estimated area under the ROC curve: the share of case and non-case pairs ranked correctly, with ties counted as half | probability from 0 to 1, with no units |
|---|
Function
Discrimination AUC as the probability of correct ranking
Maps the scores that a diagnostic test or risk prediction model gives to people with and without an outcome to a single probability: the chance that a randomly chosen person with the outcome scores higher than a randomly chosen person without it, with tied scores counted as half. S_1 and S_0 are the scores of the two people and P denotes probability. The result has no units, equals 0.5 for a score unrelated to the outcome and 1 for perfect separation, and is the area under the receiver operating characteristic (ROC) curve and the C-statistic for a binary outcome. It measures discrimination only; calibration and the costs and health effects of acting on the score are separate questions.
Computational function
Computational function: AUC from predicted risks and observed outcomes
Takes the data a validation study holds, one predicted risk or test score and one observed outcome per person, and returns the AUC. It splits the people into those with and without the outcome, compares every case with every non-case, counts the concordant and tied pairs and then applies HE-FM-AUC-001. The inputs therefore differ from the formula's variables: the formula needs pair counts, and the function builds them from patient-level data.
Inputs and outputs:
s: Vector of predicted risks or test scores, one per person, with higher values meaning greater suspicion of the outcome; required. Unit: as reported by the model, for example a probability.;y: Vector of observed outcomes in the same order, 1 for the outcome and 0 for none; required, with at least one of each. Unit: indicator.;n_1,n_0: Numbers of people with and without the outcome. Unit: count of people.;n_conc,n_tie: Concordant and tied pairs of one case and one non-case. Unit: count of pairs.;AUC: Estimated area under the ROC curve. Unit: probability from 0 to 1.Assumption: The outcome is binary and observed for everyone, with no censoring, and each person contributes one score. Pairs within the same outcome group are not compared.
Worked example (Ten-patient risk score): The article's ten patients give 21 concordant pairs and 1 tie among 24 pairs, an AUC of about 0.896.
s = [0.80,0.60,0.40,0.30,0.50,0.30,0.20,0.15,0.10,0.05]; y = [1,1,1,1,0,0,0,0,0,0]; n_1 = 4; n_0 = 6; n_conc = 21; n_tie = 1; AUC = 0.8958Worked example (Every predicted risk 0.10 higher): The shift keeps the order of the scores, so the pair counts and the AUC are unchanged, although the mean predicted risk rises from 0.34 to 0.44.
s = [0.90,0.70,0.50,0.40,0.60,0.40,0.30,0.25,0.20,0.15]; y = [1,1,1,1,0,0,0,0,0,0]; n_1 = 4; n_0 = 6; n_conc = 21; n_tie = 1; AUC = 0.8958Worked example (Score reversed): Using one minus each predicted risk reverses the order: 2 of the 24 pairs are concordant and 1 is tied, so the AUC is about 0.104, one minus the original.
s = [0.20,0.40,0.60,0.70,0.50,0.70,0.80,0.85,0.90,0.95]; y = [1,1,1,1,0,0,0,0,0,0]; n_1 = 4; n_0 = 6; n_conc = 2; n_tie = 1; AUC = 0.1042Excel:
=SUMPRODUCT((Outcome=1)*(COUNTIFS(Outcome,0,Score,"<"&Score)+0.5*COUNTIFS(Outcome,0,Score,Score)))/(COUNTIF(Outcome,1)*COUNTIF(Outcome,0))With the scores in a range named Score and the outcomes, coded 1 or 0, in a range of the same length named Outcome, the formula counts for each person with the outcome the non-cases with a lower score plus half those with the same score, and divides by the number of pairs.R:
auc_pairs <- function(s, y) { d <- outer(s[y == 1], s[y == 0], "-"); (sum(d > 0) + 0.5 * sum(d == 0)) / length(d) }The outer function forms the n_1 by n_0 matrix of score differences, so length(d) is the number of pairs. Memory grows with n_1 times n_0, so very large samples use the equivalent Wilcoxon rank form instead.Python:
def auc_pairs(s, y): cases = [a for a, o in zip(s, y) if o == 1]; controls = [b for b, o in zip(s, y) if o == 0]; return sum((a > b) + 0.5 * (a == b) for a in cases for b in controls) / (len(cases) * len(controls))Standard library only; each comparison adds as 0 or 1.Test (Pair count matches the trapezoidal ROC area): For the ten-patient data the function and the trapezoidal formula HE-FM-AUC-002 both return 21.5 divided by 24. Expected result: TRUE. Excel check:
=ABS(AUCPairs-AUCTrap)<1E-9Test (Order-preserving shift leaves the AUC unchanged): A helper range named ShiftedScore holding Score plus 0.1 returns the same AUC as the original scores. Expected result: TRUE. Excel check:
=ABS(SUMPRODUCT((Outcome=1)*(COUNTIFS(Outcome,0,ShiftedScore,"<"&ShiftedScore)+0.5*COUNTIFS(Outcome,0,ShiftedScore,ShiftedScore)))/(COUNTIF(Outcome,1)*COUNTIF(Outcome,0))-AUC)<1E-9Common error (Dividing by every pair of patients): Comparing all 45 pairs among the ten patients, rather than only the 24 pairs of one case and one non-case, gives 21.5 divided by 45, about 0.478, which wrongly suggests a score worse than chance.
Source: Hanley JA, McNeil BJ. The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology. 1982;143(1):29-36. The area equals the probability of ranking a randomly chosen diseased subject above a randomly chosen non-diseased subject, the quantity estimated by the Wilcoxon statistic.
n_1 = sum(y); n_0 = sum(1 - y); n_conc = sum_(i: y_i=1) sum_(j: y_j=0) [1(s_i > s_j)]; n_tie = sum_(i: y_i=1) sum_(j: y_j=0) [1(s_i = s_j)]; AUC = (n_conc + 0.5 * n_tie) / (n_1 * n_0)
Try this function
Implementations
Excel
AUC from pair counts in one cell
With named cells Concordant, Tied, Cases and NonCases holding n_conc, n_tie, n_1 and n_0, the formula returns the AUC. The computational function on this page counts the pairs from patient-level scores.
=(Concordant+0.5*Tied)/(Cases*NonCases)
Assumptions
Binary outcome known for every person in the AUC sample
Every person is classed as having or not having the outcome. With time-to-event data and censoring some pairs cannot be ordered, and a survival C-statistic is used instead; its versions can give different values for the same model, so the method and follow-up period are reported with it.
Higher score means higher risk in the AUC pair count
The direction of the score is fixed before counting: a pair is concordant when the person with the outcome has the higher score. If lower values indicate the condition, the scores are reversed first. Counting in the wrong direction returns one minus the AUC.
Worked examples
Ten patients with a cardiovascular risk score
In the article's illustrative example four of ten patients have an event. Of the 24 pairs of one event patient and one non-event patient, 21 are ranked correctly and 1 is tied at a predicted risk of 0.30, so the estimate is 21.5 divided by 24, about 0.896.
n_conc = 21; n_tie = 1; n_1 = 4; n_0 = 6; AUC = 0.8958
Same predicted risk for all ten patients
A model that gives all ten patients the same predicted risk ties all 24 pairs. The AUC is 0.5, the value of a score unrelated to the outcome.
n_conc = 0; n_tie = 24; n_1 = 4; n_0 = 6; AUC = 0.5
Common errors
Counting tied pairs as wins or losses in the AUC
In the ten-patient example, counting the tie at 0.30 as concordant gives 22 divided by 24, about 0.917, and dropping it gives 21 divided by 24, or 0.875. Only the half credit reproduces the area under the empirical ROC curve, about 0.896.
Reading a high AUC as evidence of calibration
The AUC depends only on the order of the scores. Adding 0.10 to every predicted risk in the ten-patient example leaves the AUC at about 0.896, but the mean predicted risk moves from 0.34 to 0.44 against an observed event rate of 0.40, and at an illustrative treatment threshold of 0.35 the number of patients flagged rises from 4 to 6. At an illustrative £400 per patient flagged, treatment costs per ten patients rise from £1,600 to £2,400. Calibration is checked separately before predicted risks enter a cost-effectiveness model.
Reporting the AUC of a score coded in the wrong direction
If a lower score indicates higher risk but the count treats higher scores as higher risk, the result is one minus the correct AUC: about 0.104 instead of 0.896 for the ten-patient example. An estimate well below 0.5 is a prompt to check the direction of the score.
Sources
Hanley and McNeil on the meaning of the ROC area
Hanley JA, McNeil BJ. The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology. 1982;143(1):29-36. Shows that the area is the probability that a randomly chosen diseased subject is ranked with greater suspicion than a randomly chosen non-diseased subject, the quantity estimated by the Wilcoxon statistic.
Steyerberg framework on the concordance statistic
Steyerberg EW, Vickers AJ, Cook NR, Gerds T, Gonen M, Obuchowski N, Pencina MJ, Kattan MW. Assessing the performance of prediction models: a framework for traditional and novel measures. Epidemiology. 2010;21(1):128-138. Section on discrimination: the c statistic is identical to the ROC area for a binary outcome, is insensitive to systematic errors in calibration as a rank order statistic, and depends on the censoring pattern when pairs that cannot be ordered are ignored.
Canonical Identity
Stable URI · Machine-readable · Resolvable · CC BY 4.0