AI Grounds

Open AI Grounds on a desktop

These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.

Guided discovery

Calibration & Reliability Diagrams

Compare confidence with observed correctness.

Your dataset

Twenty authored binary predictions form four groups of five. Every selected label is 1; known correct counts A/B/C/D are 3, 4, 5, 2. Confidence is P(1), with P(0) as its complement. Editing a group changes five probabilities while keeping selected labels and known outcomes fixed. Full accuracy stays 14/20=70%.

00252550507575100100Bin 7 (60, 70]: N=5, mean confidence 65%, observed accuracy 60%, gap 5 ppBin 8 (70, 80]: N=5, mean confidence 75%, observed accuracy 80%, gap 5 ppBin 9 (80, 90]: N=5, mean confidence 85%, observed accuracy 100%, gap 15 ppBin 10 (90, 100]: N=5, mean confidence 95%, observed accuracy 40%, gap 55 ppMean confidence (%)Observed accuracy (%)

The diagonal means equal mean confidence and observed accuracy. Vertical gaps compare those values; points use actual means, not bin midpoints. The diagram and all-example ECE always use all 20 predictions, including rejected ones.

Edit only group D’s five confidence values. Selected labels and known outcomes stay fixed. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Equal-width bins over 0..100. Exact upper boundaries belong to the bin on their left.

Keep confidence greater than or equal to the cutoff. This selects existing predictions; it does not reclassify them. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Mixed confidence · Confidences (65, 75, 85, 95) · 10 bins · Threshold 50% · Selected editor D · Full accuracy 14/20=70% · Kept 20/20

Selected group only inspects and preserves answers. Actual confidence, bin-count or threshold changes clear stale explanations. At exact 50/50 this demo consistently selects label 1; both probabilities tie.

Group evidence

All five predictions in each group select label 1. Correct count equals the number of known label-1 outcomes; the remaining outcomes have known label 0. Counts never change when confidence, bins or thresholds change.
GroupNP(0) (%)P(1) / confidence (%)CorrectAccuracy (%)BinKept?
A535653607Yes
B525754808Yes
C5158551009Yes
D559524010Yes

All confidence bins

Equal-width bins use lower-open, upper-closed intervals; the first includes zero. Mean confidence uses actual member values. Empty-bin means, accuracy and gaps are undefined, with zero contribution because count is zero. Weights use N/20.
BinInterval (%)NMean confidence (%)Observed accuracy (%)Gap (pp)Contribution (pp)
1[0, 10]0UndefinedUndefinedUndefined0
2(10, 20]0UndefinedUndefinedUndefined0
3(20, 30]0UndefinedUndefinedUndefined0
4(30, 40]0UndefinedUndefinedUndefined0
5(40, 50]0UndefinedUndefinedUndefined0
6(50, 60]0UndefinedUndefinedUndefined0
7(60, 70]5656051.25
8(70, 80]5758051.25
9(80, 90]585100153.75
10(90, 100]595405513.75

All-example ECE = Σ (N / 20) × |mean confidence − observed accuracy| = 20 pp. Contributions: 0 + 0 + 0 + 0 + 0 + 0 + 1.25 + 1.25 + 3.75 + 13.75.

pp means percentage points, not relative percent change. This is an empirical binned estimate; fewer or merged bins can hide opposing gaps without changing predictions. Zero ECE here does not establish population, subgroup or individual calibration.

All-example ECE

Count-weighted absolute bin gaps on all 20 examples.

20 pp

Σ (N / 20) × |confidence − accuracy|

Binning affects the estimate; unchanged labels can give different ECE.

Coverage

Fraction of examples kept at the confidence cutoff.

100%

20 / 20 kept

The cutoff selects a subset; it does not change the all-example diagram or correctness.

Retained accuracy

Correct label fraction among the accepted examples.

70%

14 / 20 correct

A higher confidence threshold does not guarantee higher retained accuracy.

Retained ECE: 20 pp, with accepted denominator 20. Full accuracy remains 70% on all 20.

Retained bin calculation
Equal-width bins use lower-open, upper-closed intervals; the first includes zero. Mean confidence uses actual member values. Empty-bin means, accuracy and gaps are undefined, with zero contribution because count is zero. Weights use N/20.
BinInterval (%)NMean confidence (%)Observed accuracy (%)Gap (pp)Contribution (pp)
1[0, 10]0UndefinedUndefinedUndefined0
2(10, 20]0UndefinedUndefinedUndefined0
3(20, 30]0UndefinedUndefinedUndefined0
4(30, 40]0UndefinedUndefinedUndefined0
5(40, 50]0UndefinedUndefinedUndefined0
6(50, 60]0UndefinedUndefinedUndefined0
7(60, 70]5656051.25
8(70, 80]5758051.25
9(80, 90]585100153.75
10(90, 100]595405513.75
Construction and limits

Four groups each contain five toy predictions. Known correct counts are 3, 4, 5, 2, totaling 14/20. Confidence ranges 50..100 in steps of 5, representing P(1); P(0)=100−P(1). Predicted label 1 stays fixed, including the explicit 50/50 tie convention. For this special all-label-1 dataset, positive frequency equals selected-label correctness. General top-label calibration compares selected-label confidence with selected-label correctness.

Bin count is 5 or 10 over the full 0..100 range. Intervals are (lower,upper], except [0,upper] for the first. Every case belongs to exactly one bin, including 100%. Each nonempty bin uses its actual sample mean confidence and correct-count fraction. ECE weights its absolute gap by N/20 and reports percentage points. Empty-bin confidence, accuracy and gap are undefined; zero count gives zero contribution and no plotted point.

Abstention keeps confidence≥threshold, including equality. Coverage divides accepted count by all 20. Retained accuracy divides correct accepted count by accepted count; retained ECE uses the same bin edges with accepted-count weights. With no accepted examples both retained metrics are undefined. The all-20 diagram, full accuracy and all-example ECE remain independent of the threshold.

These finite, hand-authored examples illustrate reporting and aggregation. Manually changing confidence after seeing outcomes is not a fitting or evaluation procedure. Small bins and grouping can hide errors; no confidence intervals, fitted calibrator, representative held-out evidence or population guarantee are supplied. An LLM’s stated confidence is not automatically an empirically calibrated probability; this lesson does not query an LLM or grade generated answers.

Prediction and Reset restore Mixed confidence, 10 bins, threshold 50 and editor D for the current experiment. Actual confidence/bin/threshold edits clear stale answers; inspection and unchanged edits preserve them. Free-exploration Reset begins Experiment 1.

Guo et al. · Top-label reliability bins and count-weighted ECE
scikit-learn · Empirical reliability curves and independent calibration data