AI Grounds

Open AI Grounds on a desktop

These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.

Guided discovery

Simpson's Paradox & Confounding Lab

Change task mixes; compare grouped and overall success.

Your comparison

Constructed toy counts: 100 trials per algorithm, not real data or random samples. Each trial is an easy or hard task. Changing a mix constructs new counts at fixed success rates.

Overall A 42% · B 68% · A − B = −26 pp.

0%50%100%Overall successA42%42/100B68%68/100

Algorithm A keeps 100 total trials. Hard share A is 80%. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Algorithm B keeps 100 total trials. Hard share B is 20%. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

How the rates are constructed

A: 20% × 90% + 80% × 30% = 42%.
B: 80% × 80% + 20% × 20% = 68%.

These are fixed construction rates weighted by each algorithm’s own task mix. A zero-weight term is an assumption contributing zero; it is not an observed rate for an absent group.

Easy observed: A 18/20 = 90% · B 64/80 = 80%. A − B = +10 percentage points.

Hard observed: A 24/80 = 30% · B 4/20 = 20%. A − B = +10 percentage points.

Overall gap

Difference of overall success percentages.

−26 pp

A minus B · pp = percentage points

Within-group gap

Same gap in 2 jointly represented difficulty groups.

+10 pp

A subgroup rate minus B subgroup rate

Unequal task weights can reverse or amplify the overall gap.

Confounding could occur if difficulty influences which algorithm is assigned and whether a task succeeds. Difficulty is conceptualized before evaluation here. The counts alone do not establish that causal story, rule out other confounders or prove an algorithm’s causal effect. Grouping is not always the appropriate adjustment; causal context matters.

Counts by difficulty
Current constructed counts. Each algorithm keeps 100 total trials in either view.
AlgorithmDifficultySuccessesTrialsObserved rate
A · easyeasy182090%
A · hardhard248030%
A · overalloverall4210042%
B · easyeasy648080%
B · hardhard42020%
B · overalloverall6810068%
Conventions and sources

Overall success = (easy successes + hard successes)/100. An observed subgroup rate is successes/subgroup trials and is undefined when that denominator is zero. Changing easy shares in steps of 10 keeps all constructed success counts integral. The chart always maps 0–100% to the same scale; switching views preserves counts. Both percentages and percentage-point gaps are shown exactly for this finite model.

Simpson’s reversal concerns different aggregate and within-group associations. It does not by itself determine which comparison has a causal interpretation, mean that every grouping is justified, or imply that adjusting every variable removes bias. This page supplies no causal estimator, real algorithm benchmark, significance test or automatic correction.

Berkeley · Weighted comparisons and confounding
Pearl · Understanding Simpson’s paradox and causal context