AI Grounds

Open AI Grounds on a desktop

These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.

Guided discovery

Sampling Bias Lab

Change who gets observed, then change sample size.

The synthetic population has ten equally weighted records, values 1–10, with mean 5.5. The target is always all original records. Sampling uses replacement, so records can repeat.

Ideal random benchmark: every original record has equal chance and is observed. This counterfactual comparison does not promise to force real responses or recover exited records.

Show more or fewer retained observations from the same prepared sample. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Select another prepared sample under the same method and size. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Showing sample 1 of 20, with 40 observed records. Selection frame · Ideal random.

20 sample estimates

Each numbered dot is a sample mean, not an individual record. A double ring marks the selected sample. Orange dashed reference: population mean 5.5. Blue dotted reference: collection expectation 5.5000. The two references coincide in the ideal benchmark.

SampleMean15.800025.600035.225046.050055.625065.125075.550085.625095.2750105.8250115.9250125.7000135.5750145.7500154.9000165.6250175.6500185.7750195.0250205.750015.510Observed sample mean · fixed 1–10 scale

Twenty means range from 4.9000 to 6.0500. This finite illustrative batch is not the exact sampling distribution. Increasing size keeps the same full-population axis.

Observed mean

Statistic from retained observations.

5.8000

232 / 40

Population mean

Target: all ten original records.

5.5000

(1 + 2 + … + 10) / 10

Collection expectation

Expected mean under this collection rule.

5.5000

Σ(value × keep chance) / Σ(keep chance)

Sample 1: total 232 / observed size 40 = mean 5.8000. Signed observed gap = 5.8000 − 5.5 = +0.3000. Expected bias = 5.5000 − 5.5 = +0.0000.

Simulator candidate attempts to obtain these observations: 40. Recent observed values: 4, 6, 8, 10, 6, 6, 10, 3. Observed size counts retained records. Candidates represent invitations in Nonresponse; elsewhere they are simulator proposals, not actual invitations to excluded or exited records.

Original records and observation probabilities
Original target remains all ten records. Conditional observed share = keep chance / sum of keep chances.
ValueDeclared biased mechanismActive keep chanceObserved share
1Outside frame1.00.1000
2Outside frame1.00.1000
3Outside frame1.00.1000
4Outside frame1.00.1000
5Outside frame1.00.1000
6In frame1.00.1000
7In frame1.00.1000
8In frame1.00.1000
9In frame1.00.1000
10In frame1.00.1000
Exact means of the 20 samples
Same method and size 40; selected sample labelled.
SampleTotalMeanCandidate attempts
1 · selected2325.800040
22245.600040
32095.225040
42426.050040
52255.625040
62055.125040
72225.550040
82255.625040
92115.275040
102335.825040
112375.925040
122285.700040
132235.575040
142305.750040
151964.900040
162255.625040
172265.650040
182315.775040
192015.025040
202305.750040
Assumptions and reproducible collection

The original target is a fixed bounded toy population. Every candidate independently selects one of ten records uniformly with replacement. The biased mechanism changes observation probability; every accepted record remains available for later draws. Nonresponse is newly simulated per invitation, while frame membership and survivor flags are fixed. These are illustrative mechanisms, not a causal model of why real records disappear.

A seed-1309 32-bit generator prepares 20 nonoverlapping blocks of 10000 candidate attempts, consuming two uniform positions per attempt. The first chooses a record, the second tests observation probability. Each sample shows its first n accepted records. Size edits preserve that retained prefix; changing scenario or method filters the same candidates. Finite pseudorandom blocks do not prove independence, convergence or monotone improvement.

Expected bias concerns the collection expectation minus the original target; one observed gap is finite sample error. These rules push upward by construction. Missingness need not bias a mean, and real bias can have either sign. Survivor-only sampling can suit a survivor-only target; here the target deliberately includes exited records. Ideal random is a comparison, not a practical correction promise. Weighting, standard error, intervals and identification from missing data are outside this lesson.

Berkeley SticiGui · Frame and nonresponse bias
Stambaugh · Inference about survivors