AUC estimated by counting case and non-case pairs

Compares every person with the outcome with every person without it. A pair in which the person with the outcome has the higher score counts 1, a tied pair counts one half, and the total is divided by the number of such pairs. This is the article's double sum of psi(s_i, s_j) over all n_1 times n_0 pairs, written with pair counts, and it is the quantity estimated by the Wilcoxon rank statistic.

Signature

AUC = (n_conc + 0.5 * n_tie) / (n_1 * n_0)
Inputs
InputsDefinitionUnit
n_concNumber of pairs, each of one person with the outcome and one without, in which the person with the outcome has the higher scorecount of pairs
n_tieNumber of pairs of one person with the outcome and one without in which the two scores are equalcount of pairs
n_1Number of people with the outcome (cases)count of people
n_0Number of people without the outcome (non-cases)count of people
Output
AUCEstimated area under the ROC curve: the share of case and non-case pairs ranked correctly, with ties counted as halfprobability from 0 to 1, with no units

Function

Discrimination AUC as the probability of correct ranking

Maps the scores that a diagnostic test or risk prediction model gives to people with and without an outcome to a single probability: the chance that a randomly chosen person with the outcome scores higher than a randomly chosen person without it, with tied scores counted as half. S_1 and S_0 are the scores of the two people and P denotes probability. The result has no units, equals 0.5 for a score unrelated to the outcome and 1 for perfect separation, and is the area under the receiver operating characteristic (ROC) curve and the C-statistic for a binary outcome. It measures discrimination only; calibration and the costs and health effects of acting on the score are separate questions.

Computational function

  • Computational function: AUC from predicted risks and observed outcomes

    Takes the data a validation study holds, one predicted risk or test score and one observed outcome per person, and returns the AUC. It splits the people into those with and without the outcome, compares every case with every non-case, counts the concordant and tied pairs and then applies HE-FM-AUC-001. The inputs therefore differ from the formula's variables: the formula needs pair counts, and the function builds them from patient-level data.

    Inputs and outputs: s: Vector of predicted risks or test scores, one per person, with higher values meaning greater suspicion of the outcome; required. Unit: as reported by the model, for example a probability.; y: Vector of observed outcomes in the same order, 1 for the outcome and 0 for none; required, with at least one of each. Unit: indicator.; n_1, n_0: Numbers of people with and without the outcome. Unit: count of people.; n_conc, n_tie: Concordant and tied pairs of one case and one non-case. Unit: count of pairs.; AUC: Estimated area under the ROC curve. Unit: probability from 0 to 1.

    Assumption: The outcome is binary and observed for everyone, with no censoring, and each person contributes one score. Pairs within the same outcome group are not compared.

    Worked example (Ten-patient risk score): The article's ten patients give 21 concordant pairs and 1 tie among 24 pairs, an AUC of about 0.896. s = [0.80,0.60,0.40,0.30,0.50,0.30,0.20,0.15,0.10,0.05]; y = [1,1,1,1,0,0,0,0,0,0]; n_1 = 4; n_0 = 6; n_conc = 21; n_tie = 1; AUC = 0.8958

    Worked example (Every predicted risk 0.10 higher): The shift keeps the order of the scores, so the pair counts and the AUC are unchanged, although the mean predicted risk rises from 0.34 to 0.44. s = [0.90,0.70,0.50,0.40,0.60,0.40,0.30,0.25,0.20,0.15]; y = [1,1,1,1,0,0,0,0,0,0]; n_1 = 4; n_0 = 6; n_conc = 21; n_tie = 1; AUC = 0.8958

    Worked example (Score reversed): Using one minus each predicted risk reverses the order: 2 of the 24 pairs are concordant and 1 is tied, so the AUC is about 0.104, one minus the original. s = [0.20,0.40,0.60,0.70,0.50,0.70,0.80,0.85,0.90,0.95]; y = [1,1,1,1,0,0,0,0,0,0]; n_1 = 4; n_0 = 6; n_conc = 2; n_tie = 1; AUC = 0.1042

    Excel: =SUMPRODUCT((Outcome=1)*(COUNTIFS(Outcome,0,Score,"<"&Score)+0.5*COUNTIFS(Outcome,0,Score,Score)))/(COUNTIF(Outcome,1)*COUNTIF(Outcome,0)) With the scores in a range named Score and the outcomes, coded 1 or 0, in a range of the same length named Outcome, the formula counts for each person with the outcome the non-cases with a lower score plus half those with the same score, and divides by the number of pairs.

    R: auc_pairs <- function(s, y) { d <- outer(s[y == 1], s[y == 0], "-"); (sum(d > 0) + 0.5 * sum(d == 0)) / length(d) } The outer function forms the n_1 by n_0 matrix of score differences, so length(d) is the number of pairs. Memory grows with n_1 times n_0, so very large samples use the equivalent Wilcoxon rank form instead.

    Python: def auc_pairs(s, y): cases = [a for a, o in zip(s, y) if o == 1]; controls = [b for b, o in zip(s, y) if o == 0]; return sum((a > b) + 0.5 * (a == b) for a in cases for b in controls) / (len(cases) * len(controls)) Standard library only; each comparison adds as 0 or 1.

    Test (Pair count matches the trapezoidal ROC area): For the ten-patient data the function and the trapezoidal formula HE-FM-AUC-002 both return 21.5 divided by 24. Expected result: TRUE. Excel check: =ABS(AUCPairs-AUCTrap)<1E-9

    Test (Order-preserving shift leaves the AUC unchanged): A helper range named ShiftedScore holding Score plus 0.1 returns the same AUC as the original scores. Expected result: TRUE. Excel check: =ABS(SUMPRODUCT((Outcome=1)*(COUNTIFS(Outcome,0,ShiftedScore,"<"&ShiftedScore)+0.5*COUNTIFS(Outcome,0,ShiftedScore,ShiftedScore)))/(COUNTIF(Outcome,1)*COUNTIF(Outcome,0))-AUC)<1E-9

    Common error (Dividing by every pair of patients): Comparing all 45 pairs among the ten patients, rather than only the 24 pairs of one case and one non-case, gives 21.5 divided by 45, about 0.478, which wrongly suggests a score worse than chance.

    Source: Hanley JA, McNeil BJ. The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology. 1982;143(1):29-36. The area equals the probability of ranking a randomly chosen diseased subject above a randomly chosen non-diseased subject, the quantity estimated by the Wilcoxon statistic.

    n_1 = sum(y); n_0 = sum(1 - y); n_conc = sum_(i: y_i=1) sum_(j: y_j=0) [1(s_i > s_j)]; n_tie = sum_(i: y_i=1) sum_(j: y_j=0) [1(s_i = s_j)]; AUC = (n_conc + 0.5 * n_tie) / (n_1 * n_0)

Try this function

Implementations

  • Excel

    AUC from pair counts in one cell

    With named cells Concordant, Tied, Cases and NonCases holding n_conc, n_tie, n_1 and n_0, the formula returns the AUC. The computational function on this page counts the pairs from patient-level scores.

    =(Concordant+0.5*Tied)/(Cases*NonCases)

Assumptions

  • Binary outcome known for every person in the AUC sample

    Every person is classed as having or not having the outcome. With time-to-event data and censoring some pairs cannot be ordered, and a survival C-statistic is used instead; its versions can give different values for the same model, so the method and follow-up period are reported with it.

  • Higher score means higher risk in the AUC pair count

    The direction of the score is fixed before counting: a pair is concordant when the person with the outcome has the higher score. If lower values indicate the condition, the scores are reversed first. Counting in the wrong direction returns one minus the AUC.

Worked examples

  • Ten patients with a cardiovascular risk score

    In the article's illustrative example four of ten patients have an event. Of the 24 pairs of one event patient and one non-event patient, 21 are ranked correctly and 1 is tied at a predicted risk of 0.30, so the estimate is 21.5 divided by 24, about 0.896.

    n_conc = 21; n_tie = 1; n_1 = 4; n_0 = 6; AUC = 0.8958
  • Same predicted risk for all ten patients

    A model that gives all ten patients the same predicted risk ties all 24 pairs. The AUC is 0.5, the value of a score unrelated to the outcome.

    n_conc = 0; n_tie = 24; n_1 = 4; n_0 = 6; AUC = 0.5

Common errors

  • Counting tied pairs as wins or losses in the AUC

    In the ten-patient example, counting the tie at 0.30 as concordant gives 22 divided by 24, about 0.917, and dropping it gives 21 divided by 24, or 0.875. Only the half credit reproduces the area under the empirical ROC curve, about 0.896.

  • Reading a high AUC as evidence of calibration

    The AUC depends only on the order of the scores. Adding 0.10 to every predicted risk in the ten-patient example leaves the AUC at about 0.896, but the mean predicted risk moves from 0.34 to 0.44 against an observed event rate of 0.40, and at an illustrative treatment threshold of 0.35 the number of patients flagged rises from 4 to 6. At an illustrative £400 per patient flagged, treatment costs per ten patients rise from £1,600 to £2,400. Calibration is checked separately before predicted risks enter a cost-effectiveness model.

  • Reporting the AUC of a score coded in the wrong direction

    If a lower score indicates higher risk but the count treats higher scores as higher risk, the result is one minus the correct AUC: about 0.104 instead of 0.896 for the ten-patient example. An estimate well below 0.5 is a prompt to check the direction of the score.

Sources

  • Hanley and McNeil on the meaning of the ROC area

    Hanley JA, McNeil BJ. The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology. 1982;143(1):29-36. Shows that the area is the probability that a randomly chosen diseased subject is ranked with greater suspicion than a randomly chosen non-diseased subject, the quantity estimated by the Wilcoxon statistic.

    View source →

  • Steyerberg framework on the concordance statistic

    Steyerberg EW, Vickers AJ, Cook NR, Gerds T, Gonen M, Obuchowski N, Pencina MJ, Kattan MW. Assessing the performance of prediction models: a framework for traditional and novel measures. Epidemiology. 2010;21(1):128-138. Section on discrimination: the c statistic is identical to the ROC area for a binary outcome, is insensitive to systematic errors in calibration as a rank order statistic, and depends on the censoring pattern when pairs that cannot be ordered are ignored.

    View source →

Canonical Identity

Stable URI · Machine-readable · Resolvable · CC BY 4.0