AI Grounds

Open AI Grounds on a desktop

These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.

Guided discovery

Bias-Variance Tradeoff Lab

Compare repeated fits, not just one training score.

Your dataset

Five fixed training positions X = −2, −1, 0, 1, 2. Each target is the known toy mean f(X) = 2 + X + X² plus an independent ±0.5 sign. Fit all 32 equally likely sign combinations separately. At zero noise these weighted outcomes coincide.

Probe X = 0.5 · Known mean 2.750000 · Average prediction 4.000000 · Selected prediction 3.900000 · Selected training MSE 4.240000

-404812-2-1012YX

Thin dotted = all 32 fits; wide solid indigo = selected fit; short-dashed teal = average fit; long-dashed navy = known mean. Circles = selected targets. At the probe, hollow square = selected prediction, diamond = known mean. Coincident curves do not imply extra uncertainty; the table retains all 32 outcomes.

Degree 0 is constant, 1 a line, 2 allows a quadratic, 4 interpolates the five training targets. All degrees use least squares. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Select one displayed dataset and fit from the same 32 equally weighted outcomes. Aggregate error terms stay unchanged. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Prediction location

At X = 0.5: 1.862500 expected squared error = 1.562500 squared bias + 0.050000 prediction variance + 0.250000 response noise (display rounded).

Summaries round to three decimals; formulas and tables use six. Expected error averages all 64 combinations of 32 training outcomes and two independent fresh response signs at this probe. This is an exact toy-distribution expectation, not a realized test score or a future-performance guarantee.

Squared bias

Squared miss of the average fitted prediction from the known conditional mean.

1.563

(average prediction − known mean)²

Prediction variance

Across 32 training outcomes at this fixed probe, not across X or one fit’s residuals.

0.050

Σ(prediction − average)² / 32

Response noise

Variance of a fresh ±noise response, independent of training outcomes.

0.250

amplitude² = 0.5²

Expected error

Average squared gap to a fresh response over this exact finite model.

1.863

Squared bias + variance + noise

All 32 fitted datasets
Every draw has probability 1/32. Signs correspond to ascending X positions; draw changes only the displayed selection. Targets and predictions use six-decimal display.
DrawSignsY at−2Y at−1Y at0Y at1Y at2Probe predictionTraining MSE
1− − − − −3.5000001.5000001.5000003.5000007.5000003.5000004.800000
2+ − − − −4.5000001.5000001.5000003.5000007.5000003.7000004.960000
3− + − − −3.5000002.5000001.5000003.5000007.5000003.7000004.160000
4+ + − − −4.5000002.5000001.5000003.5000007.5000003.9000004.240000
5− − + − −3.5000001.5000002.5000003.5000007.5000003.7000004.160000
6+ − + − −4.5000001.5000002.5000003.5000007.5000003.9000004.240000
7− + + − −3.5000002.5000002.5000003.5000007.5000003.9000003.440000
8+ + + − −4.5000002.5000002.5000003.5000007.5000004.1000003.440000
9− − − + −3.5000001.5000001.5000004.5000007.5000003.7000004.960000
10+ − − + −4.5000001.5000001.5000004.5000007.5000003.9000005.040000
11 (selected)− + − + −3.5000002.5000001.5000004.5000007.5000003.9000004.240000
12+ + − + −4.5000002.5000001.5000004.5000007.5000004.1000004.240000
13− − + + −3.5000001.5000002.5000004.5000007.5000003.9000004.240000
14+ − + + −4.5000001.5000002.5000004.5000007.5000004.1000004.240000
15− + + + −3.5000002.5000002.5000004.5000007.5000004.1000003.440000
16+ + + + −4.5000002.5000002.5000004.5000007.5000004.3000003.360000
17− − − − +3.5000001.5000001.5000003.5000008.5000003.7000006.560000
18+ − − − +4.5000001.5000001.5000003.5000008.5000003.9000006.640000
19− + − − +3.5000002.5000001.5000003.5000008.5000003.9000005.840000
20+ + − − +4.5000002.5000001.5000003.5000008.5000004.1000005.840000
21− − + − +3.5000001.5000002.5000003.5000008.5000003.9000005.840000
22+ − + − +4.5000001.5000002.5000003.5000008.5000004.1000005.840000
23− + + − +3.5000002.5000002.5000003.5000008.5000004.1000005.040000
24+ + + − +4.5000002.5000002.5000003.5000008.5000004.3000004.960000
25− − − + +3.5000001.5000001.5000004.5000008.5000003.9000006.640000
26+ − − + +4.5000001.5000001.5000004.5000008.5000004.1000006.640000
27− + − + +3.5000002.5000001.5000004.5000008.5000004.1000005.840000
28+ + − + +4.5000002.5000001.5000004.5000008.5000004.3000005.760000
29− − + + +3.5000001.5000002.5000004.5000008.5000004.1000005.840000
30+ − + + +4.5000001.5000002.5000004.5000008.5000004.3000005.760000
31− + + + +3.5000002.5000002.5000004.5000008.5000004.3000004.960000
32+ + + + +4.5000002.5000002.5000004.5000008.5000004.5000004.800000
Compare all degrees
Same noise and probe; average training MSE uses all 32 outcomes and five targets per fit. Squared bias need not fall at every location as degree increases.
DegreeSquared biasPrediction varianceResponse noiseExpected errorAverage training MSE
01.5625000.0500000.2500001.8625005.000000
13.0625000.0562500.2500003.3687502.950000
20.0000000.1109380.2500000.3609380.100000
30.0000000.1540040.2500000.4040040.050000
40.0000000.1851500.2500000.4351500.000000
Construction and limits

Fit polynomial coefficients to minimize the sum of squared gaps on each dataset’s five targets. The mean function is known only for constructing and evaluating this toy; it does not enter the fitting procedure. For degree 4 five distinct X positions identify a unique interpolant. Curves plot 81 evaluated X positions joined by short straight segments, including both probe positions.

Each training target’s sign is independently equally likely plus or minus. Draw ID minus1 is its five-bit mask, from lowest bit at X = −2 to highest at X = 2. A fresh response at the probe has an independent plus/minus sign and the same amplitude. Its conditional mean is f(X) and noise variance is amplitude². Average over the full32-fit distribution, using variance divisor 32. Expanding squared prediction error under this independence gives squared bias + prediction variance + response noise. Expected error is also computed directly from all 64 combinations.

A single fit’s squared residual is neither squared bias nor repeated-training prediction variance. Training draw does not change the underlying distribution. Noise-free collapses all 32 weighted outcomes to identical targets; degree 2 or higher recovers this quadratic with zero error. At X = 0.5 degree 1 has more squared bias than degree 0, so neither a monotone bias curve nor a universal U-shaped total error is assumed. No sampling estimate, confidence interval, real model-selection procedure, causal conclusion or guarantee about other populations is made.

scikit-learn · Repeated-training bias, variance and noise