AI Grounds
Open AI Grounds on a desktop
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
AI Grounds
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
Guided discovery
Move mass; distinguish surprise from uncertainty.
Four categorical outcomes A/B/C/D have probabilities summing to 100%. Entropy is their probability-weighted average surprise in bits, not accuracy. Move mass within one selected pair; the other two buckets stay fixed. The bars show probabilities on a fixed 0..100% scale.
A/B total 50%. Set A; B receives the complement. The other two buckets stay fixed. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.
Equal buckets · Masses (25, 25, 25, 25)% · Pair A/B · First 25% of 50% · Grouping AB | CD · Total 100%
Choosing a transfer pair only selects an editor and preserves answers. Mass edits change the distribution; grouping changes what the observation reports. Neither changes a measured accuracy value or trains a model.
| Outcome | Mass (%) | Probability | Surprise (bits) | Entropy contribution (bits) | Group |
|---|---|---|---|---|---|
| A | 25 | 0.250000 | 2.000000 | 0.500000 | AB |
| B | 25 | 0.250000 | 2.000000 | 0.500000 | AB |
| C | 25 | 0.250000 | 2.000000 | 0.500000 | CD |
| D | 25 | 0.250000 | 2.000000 | 0.500000 | CD |
H(X) = Σ p×surprise = 0.500000 + 0.500000 + 0.500000 + 0.500000 = 2.000000 bits.
G tells you which named group contains X. Within a positive-probability group, divide each member’s mass by the group mass. Conditional entropy averages each group’s remaining entropy using that group’s probability. It is not an unweighted sum or just one observed branch.
| Group | Group probability | Conditional outcomes | Group entropy (bits) | Weighted entropy (bits) |
|---|---|---|---|---|
| AB | 0.500000 | A: 0.500000; B: 0.500000 | 1.000000 | 0.500000 |
| CD | 0.500000 | C: 0.500000; D: 0.500000 | 1.000000 | 0.500000 |
H(X|G) = 0.500000×1.000000 + 0.500000×1.000000 = 1.000000 bits. Impossible-group terms are defined as 0; no numerical 0×Undefined product is evaluated.
Expected surprise before the group observation, in bits.
2.000000
Σx −p(x) log₂ p(x)
Four equal outcomes give the maximum 2 bits; not accuracy.
Expected remaining entropy after observing the group, in bits.
1.000000
Σg P(g) × H(X given g)
Weight possible groups; impossible-group term 0.
Expected reduction from learning G, in bits.
1.000000
2.000000 − 1.000000
At most 1 bit for this binary group observation; not always 1.
I(X;G) = H(X)−H(X|G) = 1.000000 bits. Here G is a deterministic function of X, so I(X;G)=H(G)=1.000000 bits.
All four nonnegative integer percentage masses sum 100; probabilities are masses/100. Pair editing conserves the selected pair total and fixes the other two masses. A zero-total pair cannot transfer nonexistent mass; choose another pair. Presets restore their probabilities and grouping AB | CD; the One likely bucket preset selects A/C, the others A/B. Prediction and Reset restore the current experiment; free-exploration Reset starts Experiment 1. Actual probability/grouping edits clear stale answers; pair inspection preserves mastery.
Entropy uses base 2 logarithms and full-precision arithmetic. For p>0, surprise=−log₂p and the entropy term is−p log₂p. For p=0, the outcome is impossible under the model; its surprise diverges in the limit, but its entropy contribution tends 0. These are separate statements, not numerical multiplication 0×Infinity. A certain outcome has surprise 0 and a point-mass distribution has entropy 0. Entropy is bounded by log₂4=2 bits and is not the largest single outcome surprise.
The two fixed groupings partition the four outcomes into two pairs. For each positive group, normalize its members by group mass, then compute their conditional entropy. H(X|G) is the group-probability-weighted average. A zero-probability group has no defined conditional distribution here and contributes 0 to that expected entropy. Individual branch entropy can exceed the prior; expected conditional entropy cannot.
Expected information gain is I(X;G)=H(X)−H(X|G), nonnegative and no greater than either H(X) or 1 bit here. Because this group label is a deterministic function of the outcome, gain equals H(G); this equality need not hold for general noisy observations. Tiny floating error in the nonnegative difference is clipped at 0. Editing marginal probabilities changes the model; an arbitrary before/after entropy decrease is not itself information gained from observing G.
These are authored finite distributions, not accuracy estimates, calibrated model confidences or semantic uncertainty. There is no simulation, learning, code construction, differential entropy, KL/cross-entropy, general decision tree or automatic best-observation choice. Lower entropy or greater gain alone does not prove real-world task quality.
CMU · Entropy, weighted conditional entropy and mutual information