AI Grounds
Open AI Grounds on a desktop
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
AI Grounds
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
Guided discovery
Compare observed reward with a reason to explore.
Three options give authored reward 0 or1 when tried. Each option has its own fixed outcome stream, indexed by that option’s pull count; the next reward is revealed only by a trial. This repeatable finite trace is not random sampling or evidence of an option’s true success probability. Greedy uses observed mean reward; UCB (upper confidence bound) adds a count-based exploration bonus. Policy choices use only past observations.
Mean in indigo, applied bonus in orange, read left to right. All scores share 0–6, with exact values in the table. Means stay within 0–1; UCB scores can exceed 1. Untried options have undefined means and receive separate priority before scores are compared.
c multiplies the sampled UCB exploration bonus. Greedy ignores c. Editing the rule preserves all observed trials and reward sums. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.
Pull A/B/C chooses manually. Follow policy takes the current recommendation. Both policies try each untried option first, in alphabetical order; numeric score ties also use A before B before C. Every action adds one observed reward and no other option’s observations.
6/12 trials · Total reward 2 · UCB · Strength 1 · Next option B · Decision round 7
Highest sampled score tie: B, C; choose B.
Actual history, policy or strength changes clear stale explanations and transfer answers. Unchanged controls preserve them. Scenario buttons restore their authored starting history, UCB and strength 1. The display rounds to six decimals; all calculations and choices use full precision.
Total completed pulls, including preset observations.
6/12
t=6; next decision r=t+1=7
One manual or policy action adds exactly one trial until the budget.
Sum of all observed binary outcomes.
2
1 + 0 + 1 + 0 + 0 + 0
A reward total is different from a mean or an exploration score.
Largest sampled policy score, with alphabetical ties.
B
Highest mean + c×√(2 ln(t+1)/n) after initialization
No next reward, true option quality or policy superiority is guaranteed.
n is one option’s observed trial count; mean is its reward sum divided by n. ln is the natural logarithm, a slowly growing function of decision round r=t+1. At the same round, fewer samples give a larger positive UCB bonus. The mean is an estimate from observations; the bonus is a decision term. Scores are not probabilities.
| Option | Count n | Reward sum | Mean | Applied bonus | Score | Priority? |
|---|---|---|---|---|---|---|
| A | 4 | 2 | 0.5 | 0.986385 | 1.486385 | No |
| B | 1 | 0 | 0 | 1.97277 | 1.97277 | No |
| C | 1 | 0 | 0 | 1.97277 | 1.97277 | No |
A: 2/4=0.5. UCB bonus=1×√(2 ln(7)/4)=0.986385. Score=0.5+0.986385=1.486385.
B: 0/1=0. UCB bonus=1×√(2 ln(7)/1)=1.97277. Score=0+1.97277=1.97277.
C: 0/1=0. UCB bonus=1×√(2 ln(7)/1)=1.97277. Score=0+1.97277=1.97277.
| Round | Option | Option pull | Reward | Chosen by |
|---|---|---|---|---|
| 1 | A | 1 | 1 | Preset |
| 2 | A | 2 | 0 | Preset |
| 3 | A | 3 | 1 | Preset |
| 4 | A | 4 | 0 | Preset |
| 5 | B | 1 | 0 | Preset |
| 6 | C | 1 | 0 | Preset |
This finite authored trace demonstrates allocation and decision arithmetic. It supplies no hidden success probabilities, independent identically distributed random sampling, regret estimate, confidence-coverage measurement or proof that UCB outperforms Greedy. Rewards are fixed by option-specific pull count; reset replays the same outcomes. Future outcomes never enter policy scoring.
For sampled options, default c=1 uses the UCB1 form mean+√(2 ln(r)/n), r=t+1. Other strengths are teaching variations; related algorithms can use different constants and conventions. Both policies initialize untried options first and then break exact computed score ties alphabetically. Even c=0 retains that separate initialization rule. Zero exploration bonus is not knowledge that uncertainty disappeared.
The classic stochastic regret results require bounded independent identically distributed rewards. This authored trace does not establish those assumptions or validate those results. Fewer samples enlarge a positive bonus at the same round; growing r can increase unpulled options’ bonuses. A sampled option’s bonus can fall as its own count grows. One observed win does not identify the best option or guarantee another win.
There are twelve total trials, including preset observations. At the budget, manual and policy actions are disabled; scores for hypothetical round 13 remain inspectable. There is no automatic run, random tie break, epsilon-greedy mode or selected optimal arm. Policy/strength edits preserve past observations. Reset/prediction restores Uneven trials for the current experiment; free Reset restarts Experiment 1.
Cesa-Bianchi · Bandit allocation and UCB formula