AI Grounds

Open AI Grounds on a desktop

These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.

Guided discovery

Log Loss Confidence Penalties

Same labels; different penalties.

Your dataset

Three toy binary examples have fixed known labels: A=1, B=0, C=1. Edit the probability assigned to the true class; the other class receives its complement. Log loss is −ln(pTrue), in nats. ln is the natural logarithm: more probability on what happened gives a smaller penalty. These authored predictions are not trained model outputs.

0123A0.223144B0.223144C0.223144Individual log loss (nats)

Finite bars use a fixed 0..3 nats scale. Dashed ∞ markers mean unbounded loss at exact zero true-outcome probability; they are not finite bars of length 3.

Edit only example A; the other class receives 100 minus this value. Fixed known labels stay unchanged. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

All correct · True-class percentages (80, 80, 80) · Selected editor A · Mean loss 0.223144 nats · Accuracy 3/3

Selected example only inspects and preserves answers. Actual probability changes clear stale explanations. Each probability distribution sums to 1; the three true-class probabilities across different examples need not sum to 1.

All examples

Predictions choose the larger probability. At 50/50 both labels tie; this demo selects fixed label 0 regardless of the known label. “Correct” compares the selected label with the known label. Every example contributes to mean loss.
ExampleTrue labelP(0)P(1)P(true)PredictionCorrect?Loss (nats)
A10.20.80.81Yes0.223144
B00.80.20.80Yes0.223144
C10.20.80.81Yes0.223144

Mean log loss = (0.223144 + 0.223144 + 0.223144) / 3 = 0.223144 nats. The denominator includes zero-loss, correct and incorrect examples.

Exact true-class probability 0 gives an infinite penalty; 1 gives zero. No clipping is applied. A zero probability on the other, wrong class is compatible with zero loss. Displayed decimals are rounded; the mean uses unrounded loss terms.

Mean log loss

Average negative log probability of what happened.

0.223144 nats

(0.223144 + 0.223144 + 0.223144) / 3

Lower on these examples means less penalty; it does not prove population calibration.

Accuracy

Fraction of selected labels matching the known labels.

3/3

3 correct out of 3

The same label count can hide very different probability penalties.

Largest penalty

Highest individual example loss.

0.223144 nats

Examples: A, B, C (tied)

All penalties are nonnegative; zero terms cannot cancel a positive or infinite term.

Construction and limits

Three independent true-class percentages range 0..100 in steps of 5. Each binary distribution is (p0,p1), sums to 1, and uses fixed labels A=1, B=0, C=1. For B the editor controls p0, not p1. The predicted label is 1 if p1 exceeds 0.5; otherwise it is 0. Both maxima are exposed at an exact tie.

Each penalty is −ln(pTrue), using natural logarithms and units called nats. For 0<pTrue<1 the value is positive; at 1 it is zero. As pTrue approaches 0 the penalty grows without bound; at exact 0 its extended value is ∞. Mean loss sums all three penalties and divides by 3. Finite supported probabilities are at least 0.05, giving a largest finite loss about 2.995732; the fixed chart axis 0..3 is not a cap on infinite loss.

This lesson shows the mathematical endpoints without epsilon clipping. Numerical libraries may clip exact endpoints for finite computations. These toy predictions illustrate scoring after known outcomes, not a procedure for choosing predictions after observing held-out labels. No training, calibration frequencies, population risk estimate or universal good-loss threshold is supplied. Higher confidence alone is not always better: a more certain wrong prediction is penalized more.

Prediction and Reset restore the current experiment baseline and selected example. Actual probability edits clear stale answers; inspection and unchanged edits preserve them. Free-exploration Reset begins Experiment 1. The nearby Cross Entropy Loss Explorer covers binary, categorical and multi-label target shapes; this lesson focuses on heterogeneous penalties and averages at the same accuracy.

scikit-learn · Log-loss formula and averaging
scikit-learn · Natural logarithm and numerical endpoint clipping