Signature
AUC_trap = sum_(k=1)^K [(FPR_end_k - FPR_start_k) * (TPR_start_k + TPR_end_k) / 2]
| Inputs | Definition | Unit |
|---|---|---|
FPR_end_k | False positive rate, one minus specificity, at the end of segment k | proportion from 0 to 1 |
FPR_start_k | False positive rate, one minus specificity, at the start of segment k | proportion from 0 to 1 |
TPR_start_k | Sensitivity (true positive rate) at the start of segment k | proportion from 0 to 1 |
TPR_end_k | Sensitivity (true positive rate) at the end of segment k | proportion from 0 to 1 |
AUC_trap | Area under the empirical ROC curve from the trapezoidal rule | probability from 0 to 1, with no units |
|---|
KNumber of segments between consecutive points of the empirical ROC curve, which runs from (0, 0) to (1, 1) (count)
Function
Discrimination AUC as the probability of correct ranking
Maps the scores that a diagnostic test or risk prediction model gives to people with and without an outcome to a single probability: the chance that a randomly chosen person with the outcome scores higher than a randomly chosen person without it, with tied scores counted as half. S_1 and S_0 are the scores of the two people and P denotes probability. The result has no units, equals 0.5 for a score unrelated to the outcome and 1 for perfect separation, and is the area under the receiver operating characteristic (ROC) curve and the C-statistic for a binary outcome. It measures discrimination only; calibration and the costs and health effects of acting on the score are separate questions.
Try this function
Implementations
Excel
Trapezoidal ROC area from segment ranges
With the K segments in four named ranges of equal length, FPRStart, FPREnd, TPRStart and TPREnd, SUMPRODUCT adds each width multiplied by the mean height.
=SUMPRODUCT(FPREnd-FPRStart,(TPRStart+TPREnd)/2)
Assumptions
ROC points at every distinct threshold in order
The points are taken at every distinct score, sorted by increasing false positive rate, from (0, 0) to (1, 1), so each segment starts where the previous one ends. When a case and a non-case share a score, the threshold passes both at once and the segment through them is diagonal.
Worked examples
Ten-patient ROC curve in five segments
The article's curve runs through (0, 0), (0, 0.5), (1/6, 0.5), (1/6, 0.75), (2/6, 1) and (1, 1). The two vertical segments add nothing and the other three add about 0.0833, 0.1458 and 0.6667, a total of about 0.896, the same as the pair count.
K = 5; FPR_start_k = [0,0,0.16667,0.16667,0.33333]; FPR_end_k = [0,0.16667,0.16667,0.33333,1]; TPR_start_k = [0,0.5,0.5,0.75,1]; TPR_end_k = [0.5,0.5,0.75,1,1]; AUC_trap = 0.8958
Diagonal ROC curve of a score unrelated to the outcome
A single segment from (0, 0) to (1, 1), the line of random guessing, encloses an area of 0.5.
K = 1; FPR_start_k = [0]; FPR_end_k = [1]; TPR_start_k = [0]; TPR_end_k = [1]; AUC_trap = 0.5
Common errors
Using rectangles instead of trapezoids through tied scores
Drawing the tie as a step gives a rectangle in place of the diagonal segment. Moving across before up gives 0.875 for the ten-patient curve, and moving up before across gives about 0.917, against about 0.896 from the trapezoid. These are the pessimistic and optimistic segments that Fawcett describes for equally scored cases.
Sources
Fawcett trapezoidal algorithm for the ROC area
Fawcett T. An introduction to ROC analysis. Pattern Recognition Letters. 2006;27(8):861-874. Section 7 and Algorithm 2 (the area added as successive trapezoids, each its base multiplied by its mean height) and Fig. 6 (optimistic, pessimistic and expected segments for equally scored instances).
Canonical Identity
Stable URI · Machine-readable · Resolvable · CC BY 4.0