AI Grounds

Open AI Grounds on a desktop

These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.

Guided discovery

Entropy & Information Starter

Move mass; distinguish surprise from uncertainty.

Your dataset

Four categorical outcomes A/B/C/D have probabilities summing to 100%. Entropy is their probability-weighted average surprise in bits, not accuracy. Move mass within one selected pair; the other two buckets stay fixed. The bars show probabilities on a fixed 0..100% scale.

0%50%100%A25%B25%C25%D25%Probability (fixed 0..100%)

A/B total 50%. Set A; B receives the complement. The other two buckets stay fixed. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Observed grouping

Equal buckets · Masses (25, 25, 25, 25)% · Pair A/B · First 25% of 50% · Grouping AB | CD · Total 100%

Choosing a transfer pair only selects an editor and preserves answers. Mass edits change the distribution; grouping changes what the observation reports. Neither changes a measured accuracy value or trains a model.

All outcomes (X)

For p>0, surprise=−log₂(p); entropy contribution=p×surprise. At p=0, surprise has an infinite limit and the outcome is impossible under this model; the entropy term is 0 by continuity. Base 2 units are bits; probabilities and expectations are not observed correctness.
OutcomeMass (%)ProbabilitySurprise (bits)Entropy contribution (bits)Group
A250.2500002.0000000.500000AB
B250.2500002.0000000.500000AB
C250.2500002.0000000.500000CD
D250.2500002.0000000.500000CD

H(X) = Σ p×surprise = 0.500000 + 0.500000 + 0.500000 + 0.500000 = 2.000000 bits.

Groups (G)

G tells you which named group contains X. Within a positive-probability group, divide each member’s mass by the group mass. Conditional entropy averages each group’s remaining entropy using that group’s probability. It is not an unweighted sum or just one observed branch.

Conditional probabilities sum 1 within each possible group. A zero-probability group has Undefined conditional distribution/entropy and weighted term 0. One rare branch can have more entropy than the prior; the probability-weighted average cannot.
GroupGroup probabilityConditional outcomesGroup entropy (bits)Weighted entropy (bits)
AB0.500000A: 0.500000; B: 0.5000001.0000000.500000
CD0.500000C: 0.500000; D: 0.5000001.0000000.500000

H(X|G) = 0.500000×1.000000 + 0.500000×1.000000 = 1.000000 bits. Impossible-group terms are defined as 0; no numerical 0×Undefined product is evaluated.

H(X)

Expected surprise before the group observation, in bits.

2.000000

Σx −p(x) log₂ p(x)

Four equal outcomes give the maximum 2 bits; not accuracy.

H(X|G)

Expected remaining entropy after observing the group, in bits.

1.000000

Σg P(g) × H(X given g)

Weight possible groups; impossible-group term 0.

Information gain

Expected reduction from learning G, in bits.

1.000000

2.000000 − 1.000000

At most 1 bit for this binary group observation; not always 1.

I(X;G) = H(X)−H(X|G) = 1.000000 bits. Here G is a deterministic function of X, so I(X;G)=H(G)=1.000000 bits.

Construction and limits

All four nonnegative integer percentage masses sum 100; probabilities are masses/100. Pair editing conserves the selected pair total and fixes the other two masses. A zero-total pair cannot transfer nonexistent mass; choose another pair. Presets restore their probabilities and grouping AB | CD; the One likely bucket preset selects A/C, the others A/B. Prediction and Reset restore the current experiment; free-exploration Reset starts Experiment 1. Actual probability/grouping edits clear stale answers; pair inspection preserves mastery.

Entropy uses base 2 logarithms and full-precision arithmetic. For p>0, surprise=−log₂p and the entropy term is−p log₂p. For p=0, the outcome is impossible under the model; its surprise diverges in the limit, but its entropy contribution tends 0. These are separate statements, not numerical multiplication 0×Infinity. A certain outcome has surprise 0 and a point-mass distribution has entropy 0. Entropy is bounded by log₂4=2 bits and is not the largest single outcome surprise.

The two fixed groupings partition the four outcomes into two pairs. For each positive group, normalize its members by group mass, then compute their conditional entropy. H(X|G) is the group-probability-weighted average. A zero-probability group has no defined conditional distribution here and contributes 0 to that expected entropy. Individual branch entropy can exceed the prior; expected conditional entropy cannot.

Expected information gain is I(X;G)=H(X)−H(X|G), nonnegative and no greater than either H(X) or 1 bit here. Because this group label is a deterministic function of the outcome, gain equals H(G); this equality need not hold for general noisy observations. Tiny floating error in the nonnegative difference is clipped at 0. Editing marginal probabilities changes the model; an arbitrary before/after entropy decrease is not itself information gained from observing G.

These are authored finite distributions, not accuracy estimates, calibrated model confidences or semantic uncertainty. There is no simulation, learning, code construction, differential entropy, KL/cross-entropy, general decision tree or automatic best-observation choice. Lower entropy or greater gain alone does not prove real-world task quality.

CMU · Entropy, weighted conditional entropy and mutual information
SciPy · Entropy and logarithm-base units