AI Grounds

Open AI Grounds on a desktop

These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.

Guided discovery

ROC, AUC & Thresholds Lab

Move one operating point; inspect the whole ranking.

Your dataset

Twelve fixed scored cases: Cases 1..6 actually positive, Cases 7..12 negative. Predict Positive when score ≥ threshold, including equality. Scores are ranking signals, not calibrated probabilities. Scenario changes preserve the threshold and actual labels while changing scores.

Threshold 0.85 · TP 1 · FP 0 · FN 5 · TN 6 · TPR 1/6 · FPR 0/6 · AUC 0.833333

0.00.00.50.51.01.0True-positive rateFalse-positive rate

Circles = grouped score-sweep vertices. Hollow diamond = selected operating point. Pale fill = area under the whole ROC curve. Dashed diagonal = tied-ranking reference (AUC 0.5). Connecting segments summarize area; tied groups may leave intermediate points unavailable to a hard threshold.

Score ≥ threshold predicts Positive. Move in 0.05 steps. Equal scores move together; intervals with no scores can preserve the same point. This changes predictions, not scores or AUC. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

AUC = (29 wins + 0.5 × 2 ties)/36 = 30/36 = 0.833333. Whole-ranking area is independent of the selected threshold.

False-positive rate

Accepted actual negatives among all six actual negatives.

0.0%

FPR = FP/(FP + TN) = 0/6

True-positive rate

Found actual positives among all six actual positives; this is recall.

16.7%

TPR = TP/(TP + FN) = 1/6

AUC

Area under the entire ROC curve; a ranking statistic from 0 to 1.

0.833

(29 wins + 0.5 × 2 ties)/36

Neither current-threshold accuracy nor calibration.

Twelve scores and decisions
Actual labels and IDs stay fixed. Current score ≥ 0.85 predicts Positive; every case contributes to one count.
CaseScoreActualPredictedConfusion cell
Case 10.90PositivePositiveTP
Case 20.80PositiveNegativeFN
Case 30.70PositiveNegativeFN
Case 40.70PositiveNegativeFN
Case 50.60PositiveNegativeFN
Case 60.40PositiveNegativeFN
Case 70.80NegativeNegativeTN
Case 80.50NegativeNegativeTN
Case 90.40NegativeNegativeTN
Case 100.30NegativeNegativeTN
Case 110.20NegativeNegativeTN
Case 120.10NegativeNegativeTN
ROC vertices and area
Representative distinct-score cutoffs plus 1.00 (above all scores here). Equal scores enter together. Each segment’s area is its FPR width × average endpoint TPR; vertical width contributes zero.
ThresholdFPTPFPRTPRArea from previous
1.00000.0000000.0000000.000000
0.90010.0000000.1666670.000000
0.80120.1666670.3333330.041667
0.70140.1666670.6666670.000000
0.60150.1666670.8333330.000000
0.50250.3333330.8333330.138889
0.40360.5000001.0000000.152778
0.30460.6666671.0000000.166667
0.20560.8333331.0000000.166667
0.10661.0000001.0000000.166667

Sum of full-precision segment areas = 0.833333. Displayed terms are rounded.

Construction and limits

Mostly ordered scores for Cases 1..12 are 0.90,0.80,0.70,0.70,0.60,0.40,0.80,0.50,0.40,0.30,0.20,0.10. Reversed uses 1 − score with the same actual labels; All tied uses 0.50. Scores and threshold are compared on integer 0.05 ticks. Every dataset has six actual positives and six negatives, so no zero-denominator ROC case is introduced.

ROC uses actual-class denominators. It does not plot precision, accuracy or predicted-positive proportions. Sweep score groups in descending order; lowering the inclusive threshold can increase or leave unchanged each rate. Tie groups jump together, so a straight area segment does not make every intermediate point a separate hard-threshold outcome.

AUC is the trapezoidal area and equals (positive-over-negative pair wins+half ties)/36. Here each of the 36 pairs chooses one positive and one negative uniformly from this finite sample. Mostly ordered has 29 wins and 2 ties, Reversed 5 wins and 2 ties, All tied 0 wins and 36 ties. AUC 0.5 need not mean random score generation. AUC does not select an optimal threshold, measure calibrated probabilities, encode error costs or guarantee future performance. No fitting, ROC uncertainty, class-prevalence editing or precision–recall curve is implemented.

scikit-learn · ROC thresholds and score grouping
scikit-learn · Trapezoidal area
mlr3measures · Ranking pairs and tied scores
Google · ROC and AUC interpretation