AI Grounds

Open AI Grounds on a desktop

These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.

Guided discovery

Precision-Recall Curves & Imbalance

Make positives rare; inspect accepted predictions.

Your dataset

Six positive cases and six negative score types. Copy every negative score equally to change prevalence (the actual-positive fraction) while preserving each class’s score frequencies. Copies are illustrative accounting, not new independent observations. Scores are fixed ranking signals; score ≥ threshold predicts Positive, including equality.

6 positives · 6 negatives · prevalence 50.0% · threshold 0.55 · TP 5 · FP 1 · FN 1 · TN 5 · precision 0.833333 · recall 0.833333 · AP 0.794444

0.00.00.50.51.01.0PrecisionRecall

Circles = grouped threshold outcomes. Hollow diamond = current operating point. Pale rectangles = average precision (AP) weights. Dashed line = prevalence, the precision when all cases are predicted positive. Open square at (recall 0, precision 1) = drawing convention with no threshold. Steps summarize AP; intermediate points need not be attainable hard cutoffs.

Copy all six negative score types equally. Six positives remain fixed; negatives = 6 × copies. This changes prevalence and precision, not the positive scores or recall at a fixed threshold. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Score ≥ threshold predicts Positive. Equal scores enter together. Lowering the cutoff can raise, lower or preserve precision. A cutoff edit selects one point without changing AP. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

AP = sum(recall increase × score-group precision) = 0.794444. Non-interpolated recall weights; not trapezoidal PR area, ROC AUC or one cutoff’s precision.

Precision

True positives among accepted predictions.

83.3%

TP/(TP + FP) = 5/6

Uses predicted positives, not all cases.

Recall

Found actual positives among all six actual positives.

83.3%

TP/(TP + FN) = 5/6

Changing uniform negative copies preserves this rate.

Average precision

Whole-curve summary from each score group’s precision weighted by its increase in recall.

0.794

AP = sum(Δrecall × precision)

Depends on class balance; unchanged by one cutoff edit.

Cases and decisions
P1..P6 are fixed positives. N-type.copy identifies each illustrative negative copy. Score ≥ 0.55 predicts Positive.
CaseScoreActualPredictedConfusion cell
P10.90PositivePositiveTP
P20.80PositivePositiveTP
P30.70PositivePositiveTP
P40.70PositivePositiveTP
P50.60PositivePositiveTP
P60.40PositiveNegativeFN
N1.10.80NegativePositiveFP
N2.10.50NegativeNegativeTN
N3.10.40NegativeNegativeTN
N4.10.30NegativeNegativeTN
N5.10.20NegativeNegativeTN
N6.10.10NegativeNegativeTN
Score groups and AP
Distinct scores enter together in descending order. AP term = increase in recall × this group’s precision; zero recall-width contributes zero. The drawing endpoint has no threshold and no AP term.
ThresholdTPFPRecallPrecisionΔrecallAP term
0.90100.1666671.0000000.1666670.166667
0.80210.3333330.6666670.1666670.111111
0.70410.6666670.8000000.3333330.266667
0.60510.8333330.8333330.1666670.138889
0.50520.8333330.7142860.0000000.000000
0.40631.0000000.6666670.1666670.111111
0.30641.0000000.6000000.0000000.000000
0.20651.0000000.5454550.0000000.000000
0.10661.0000000.5000000.0000000.000000

Sum of full-precision AP terms = 0.794444. Displayed terms are rounded.

Construction and limits

Mostly ordered positive scores: 0.90, 0.80, 0.70, 0.70, 0.60, 0.40. Negative score types: 0.80, 0.50, 0.40, 0.30, 0.20, 0.10. Reversed uses 1 − score; All tied uses 0.50. Negative copies per score repeats each negative type equally, preserving both classes’ score distributions. Prevalence = 6/(6 + 6 × copies), from 50% at one copy to 20% at four. Scores and threshold compare on integer 0.05 ticks.

Precision = TP/(TP + FP); recall = TP/6. With no accepted cases, precision is undefined. The open drawing square (0,1) has no threshold; it is not an observed precision of 1. If all cases are accepted, precision equals prevalence. Constant tied scores give one grouped recall jump and AP exactly equal to prevalence here, without proving random score generation.

Descending score groups define non-interpolated AP = sum(Δrecall × current-group precision). Each rectangle uses the precision after that entire tied group enters. This differs from trapezoidal PR area, ROC AUC and an unweighted average of precision values. Recall cannot fall as a cutoff lowers, but precision can move either way. Step interiors and the drawing endpoint are not additional hard-threshold outcomes.

No model training, calibration, error-cost policy, automatic optimal threshold, ROC plot, uncertainty or real independent sampling is implemented. Both actual classes are always present. Replication isolates class balance without enlarging the independent evidence. Representative evaluation data and task costs are needed beyond these finite values.

scikit-learn · PR thresholds and drawing endpoint
scikit-learn · Non-interpolated average precision
scikit-learn · Precision-recall interpretation