AI Grounds
Open AI Grounds on a desktop
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
AI Grounds
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
Guided discovery
Change the budget. Separate finding an answer from selecting one.
Eight authored multiple-choice items have disclosed gold labels. Three fixed authored sample batches illustrate sampled answer outcomes; no language model, random generator or code evaluator runs here. A/B/C are answer labels, not confidence levels. ✓ means a gold match; × means a different answer. Only the first Samples n columns are counted. The other columns remain visible as Unused.
Q7/Q8 have disclosed authored exposure marks and deliberately correct sample outcomes. These marks illustrate a filtering effect; they do not detect real training contamination, estimate its causal effect or prove unmarked items are clean. Omitted rows stay visible for comparison and leave every aggregate denominator.
| Item and question | Gold | Sample 1 | Sample 2 | Sample 3 | Sample 4 | Sample 5 | Sample 6 |
|---|---|---|---|---|---|---|---|
| Q1 · 2 + 3?A=5, B=6, C=7Included · Unmarked | A | ✓ A | × B | ✓ A | × C | ✓ A | × B |
| Q2 · 3 × 2?A=5, B=6, C=9Included · Unmarked | B | × A | ✓ B | × A | × A | × C | ✓ B |
| Q3 · 9 − 4?A=4, B=6, C=5Included · Unmarked | C | × A | × B | × A | × B | × A | × B |
| Q4 · 8 ÷ 2?A=4, B=2, C=6Included · Unmarked | A | × B | ✓ A | × C | × B | × C | ✓ A |
| Q5 · 2²?A=2, B=4, C=8Included · Unmarked | B | ✓ B | × A | ✓ B | ✓ B | × C | ✓ B |
| Q6 · 7 − 2?A=3, B=4, C=5Included · Unmarked | C | ✓ C | × A | ✓ C | × B | ✓ C | × A |
| Q7 · 1 + 1?A=2, B=1, C=3Included · Marked exposure | A | ✓ A | ✓ A | ✓ A | ✓ A | ✓ A | ✓ A |
| Q8 · 3 + 1?A=2, B=4, C=3Included · Marked exposure | B | ✓ B | ✓ B | ✓ B | ✓ B | ✓ B | ✓ B |
Count the first n stored answers per item. Shrinking n clamps k down if necessary; raising n keeps the current k. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.
Size of the subset whose at-least-one-correct event is evaluated. Must be ≤n. This does not change answers or how many answers voting uses. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.
8 included items · n=6, k=1. All-sample accuracy = 26/48 = 0.541667. 2 included tied votes abstain and count incorrect. First attempt and voting do not change when only k changes.
Every actual control or preset change clears stale explanation/transfer answers; reapplying the same state preserves them. Sample batch chooses stored outcomes. It does not regenerate or train a model. Vote uses all n active columns, not only k.
Gold matches in Sample 1 over included items.
0.625
first correct / included item count
5/8. The first sampled answer is not a greedy-decoding claim.
Mean per-item at-least-one-correct subset probability.
0.541667
mean(1 − C(n−c,k) / C(n,k))
k≤n. Requires gold/evaluator knowledge; it does not select the correct answer for deployment.
Unique most frequent answer matches over included items.
0.625
correct plurality votes / included item count
5/8; ties abstain and count incorrect. A unique plurality need not exceed half the samples.
C(n,k) counts the ways to choose k samples without regard to order. Of those subsets, C(n−c,k) contain no gold match. Subtract their ratio from 1. C(m,k)=0 when m<k. An item with no correct samples has pass 0; when k=n, an item with any correct sample has pass 1. The table reports every item; only Included rows enter the mean.
Estimated pass@k = mean over included items of [1 − C(n−c,k)/C(n,k)]
| Item / inclusion | Correct c | Failure subsets | All subsets | Estimated pass | A / B / C counts | Vote | Gold match |
|---|---|---|---|---|---|---|---|
| Q1 · Included | 3/6 | 3 | 6 | 0.5 | 3 / 2 / 1 | A | Yes |
| Q2 · Included | 2/6 | 4 | 6 | 0.333333 | 3 / 2 / 1 | A | No |
| Q3 · Included | 0/6 | 6 | 6 | 0 | 3 / 3 / 0 | Tie · Abstain | No |
| Q4 · Included | 2/6 | 4 | 6 | 0.333333 | 2 / 2 / 2 | Tie · Abstain | No |
| Q5 · Included | 4/6 | 2 | 6 | 0.666667 | 1 / 4 / 1 | B | Yes |
| Q6 · Included | 3/6 | 3 | 6 | 0.5 | 2 / 1 / 3 | C | Yes |
| Q7 · Included | 6/6 | 0 | 6 | 1 | 6 / 0 / 0 | A | Yes |
| Q8 · Included | 6/6 | 0 | 6 | 1 | 0 / 6 / 0 | B | Yes |
Same-setting all-eight reference: first=0.625, sample=0.541667, pass@1=0.541667, vote=0.625. Current aggregate uses 8 items. Filtering changes the evaluated set, not any answer.
All three rows use the current n, k and exposure policy. They describe fixed authored outcome batches, not new independent model runs. Changing the batch changes the observed answers; it does not change the model or establish that its abilities improved.
| Batch | First-attempt accuracy | All-sample accuracy / pass@1 | Estimated pass@1 | Vote accuracy |
|---|---|---|---|---|
| A · Current | 0.625 | 0.541667 | 0.541667 | 0.625 |
| B | 0.625 | 0.5625 | 0.5625 | 0.625 |
| C | 0.875 | 0.583333 | 0.583333 | 0.75 |
Pass@1 across A/B/C: mean 0.5625 · population variance 0.000289 · standard deviation 0.01701 · range 0.541667 to 0.583333.
variance = sum((batch pass − mean)²)/3; standard deviation = sqrt(variance)
These descriptive values use fractions, not percentage points; variance has squared-fraction units. They are not a standard error, confidence interval, significance test or guarantee about future performance. Three hand-authored batches do not provide representative uncertainty estimates.
The finite model has three fixed authored batches, eight multiple-choice items, integer n=1..6, integer k=1..n and two exposure policies. Gold labels and options are disclosed. There is no model, random sampling, code generation, execution, evaluator error, training, temperature or contamination detector. Each batch is a stored six-answer list per item; current n takes its prefix. Shrinking n clamps k down, while increasing n does not restore a larger k. Comparisons should state their item set, budget and outcome batch.
First-attempt accuracy is exact gold matches in Sample 1 divided by included item count. It is not greedy accuracy. All-sample accuracy is total gold matches divided by included items times n. Per-item pass@k=1−C(n−c,k)/C(n,k), averaged equally over included items. It is the exact probability that a uniformly chosen size-k subset of the observed samples contains at least one correct answer, and the standard unbiased estimator for fresh-attempt pass@k under the usual independent identically distributed sampling assumptions. This authored demonstration does not validate those assumptions. Do not replace the finite-sample estimator with 1−(1−c/n)^k or treat larger k as model improvement. k>n is not admitted. Equal n makes the mean pass@1 equal all-sample accuracy.
Voting uses answer-label frequencies over all n active samples. This lab’s majority-vote example means unique plurality: the largest count wins, even if below half. Equal largest counts abstain and count incorrect; ties remain in the denominator. It does not use gold to select an answer. An at-least-one-correct metric uses oracle/evaluator correctness after sampling; it does not supply an answer-selection policy. Pass@k, first/sample accuracy and voting can disagree, and adding observed samples need not monotonically raise a fixed-k estimate.
Q7/Q8 are deliberately marked exposure examples with correct outcomes, used to construct an evaluation-subset effect. Omitting them leaves six items and changes only aggregation. The same-setting all-eight reference keeps current batch/n/k. This difference is not a measured causal training-contamination penalty; omission does not detect all exposure or prove remaining items are clean. Cross-batch mean, population variance (divide by 3), standard deviation and range describe only these authored values. No standard error, confidence interval, significance claim or real benchmark leaderboard is supplied. Presets restore their complete states; Reset/prediction changes restore One chance for the current experiment, while free Reset starts Experiment 1.
HumanEval · Per-item pass@k estimator and averaging