AI Grounds
Open AI Grounds on a desktop
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
AI Grounds
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
Guided discovery
Compare repeated fits, not just one training score.
Five fixed training positions X = −2, −1, 0, 1, 2. Each target is the known toy mean f(X) = 2 + X + X² plus an independent ±0.5 sign. Fit all 32 equally likely sign combinations separately. At zero noise these weighted outcomes coincide.
Probe X = 0.5 · Known mean 2.750000 · Average prediction 4.000000 · Selected prediction 3.900000 · Selected training MSE 4.240000
Thin dotted = all 32 fits; wide solid indigo = selected fit; short-dashed teal = average fit; long-dashed navy = known mean. Circles = selected targets. At the probe, hollow square = selected prediction, diamond = known mean. Coincident curves do not imply extra uncertainty; the table retains all 32 outcomes.
Degree 0 is constant, 1 a line, 2 allows a quadratic, 4 interpolates the five training targets. All degrees use least squares. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.
Select one displayed dataset and fit from the same 32 equally weighted outcomes. Aggregate error terms stay unchanged. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.
At X = 0.5: 1.862500 expected squared error = 1.562500 squared bias + 0.050000 prediction variance + 0.250000 response noise (display rounded).
Summaries round to three decimals; formulas and tables use six. Expected error averages all 64 combinations of 32 training outcomes and two independent fresh response signs at this probe. This is an exact toy-distribution expectation, not a realized test score or a future-performance guarantee.
Squared miss of the average fitted prediction from the known conditional mean.
1.563
(average prediction − known mean)²
Across 32 training outcomes at this fixed probe, not across X or one fit’s residuals.
0.050
Σ(prediction − average)² / 32
Variance of a fresh ±noise response, independent of training outcomes.
0.250
amplitude² = 0.5²
Average squared gap to a fresh response over this exact finite model.
1.863
Squared bias + variance + noise
| Draw | Signs | Y at−2 | Y at−1 | Y at0 | Y at1 | Y at2 | Probe prediction | Training MSE |
|---|---|---|---|---|---|---|---|---|
| 1 | − − − − − | 3.500000 | 1.500000 | 1.500000 | 3.500000 | 7.500000 | 3.500000 | 4.800000 |
| 2 | + − − − − | 4.500000 | 1.500000 | 1.500000 | 3.500000 | 7.500000 | 3.700000 | 4.960000 |
| 3 | − + − − − | 3.500000 | 2.500000 | 1.500000 | 3.500000 | 7.500000 | 3.700000 | 4.160000 |
| 4 | + + − − − | 4.500000 | 2.500000 | 1.500000 | 3.500000 | 7.500000 | 3.900000 | 4.240000 |
| 5 | − − + − − | 3.500000 | 1.500000 | 2.500000 | 3.500000 | 7.500000 | 3.700000 | 4.160000 |
| 6 | + − + − − | 4.500000 | 1.500000 | 2.500000 | 3.500000 | 7.500000 | 3.900000 | 4.240000 |
| 7 | − + + − − | 3.500000 | 2.500000 | 2.500000 | 3.500000 | 7.500000 | 3.900000 | 3.440000 |
| 8 | + + + − − | 4.500000 | 2.500000 | 2.500000 | 3.500000 | 7.500000 | 4.100000 | 3.440000 |
| 9 | − − − + − | 3.500000 | 1.500000 | 1.500000 | 4.500000 | 7.500000 | 3.700000 | 4.960000 |
| 10 | + − − + − | 4.500000 | 1.500000 | 1.500000 | 4.500000 | 7.500000 | 3.900000 | 5.040000 |
| 11 (selected) | − + − + − | 3.500000 | 2.500000 | 1.500000 | 4.500000 | 7.500000 | 3.900000 | 4.240000 |
| 12 | + + − + − | 4.500000 | 2.500000 | 1.500000 | 4.500000 | 7.500000 | 4.100000 | 4.240000 |
| 13 | − − + + − | 3.500000 | 1.500000 | 2.500000 | 4.500000 | 7.500000 | 3.900000 | 4.240000 |
| 14 | + − + + − | 4.500000 | 1.500000 | 2.500000 | 4.500000 | 7.500000 | 4.100000 | 4.240000 |
| 15 | − + + + − | 3.500000 | 2.500000 | 2.500000 | 4.500000 | 7.500000 | 4.100000 | 3.440000 |
| 16 | + + + + − | 4.500000 | 2.500000 | 2.500000 | 4.500000 | 7.500000 | 4.300000 | 3.360000 |
| 17 | − − − − + | 3.500000 | 1.500000 | 1.500000 | 3.500000 | 8.500000 | 3.700000 | 6.560000 |
| 18 | + − − − + | 4.500000 | 1.500000 | 1.500000 | 3.500000 | 8.500000 | 3.900000 | 6.640000 |
| 19 | − + − − + | 3.500000 | 2.500000 | 1.500000 | 3.500000 | 8.500000 | 3.900000 | 5.840000 |
| 20 | + + − − + | 4.500000 | 2.500000 | 1.500000 | 3.500000 | 8.500000 | 4.100000 | 5.840000 |
| 21 | − − + − + | 3.500000 | 1.500000 | 2.500000 | 3.500000 | 8.500000 | 3.900000 | 5.840000 |
| 22 | + − + − + | 4.500000 | 1.500000 | 2.500000 | 3.500000 | 8.500000 | 4.100000 | 5.840000 |
| 23 | − + + − + | 3.500000 | 2.500000 | 2.500000 | 3.500000 | 8.500000 | 4.100000 | 5.040000 |
| 24 | + + + − + | 4.500000 | 2.500000 | 2.500000 | 3.500000 | 8.500000 | 4.300000 | 4.960000 |
| 25 | − − − + + | 3.500000 | 1.500000 | 1.500000 | 4.500000 | 8.500000 | 3.900000 | 6.640000 |
| 26 | + − − + + | 4.500000 | 1.500000 | 1.500000 | 4.500000 | 8.500000 | 4.100000 | 6.640000 |
| 27 | − + − + + | 3.500000 | 2.500000 | 1.500000 | 4.500000 | 8.500000 | 4.100000 | 5.840000 |
| 28 | + + − + + | 4.500000 | 2.500000 | 1.500000 | 4.500000 | 8.500000 | 4.300000 | 5.760000 |
| 29 | − − + + + | 3.500000 | 1.500000 | 2.500000 | 4.500000 | 8.500000 | 4.100000 | 5.840000 |
| 30 | + − + + + | 4.500000 | 1.500000 | 2.500000 | 4.500000 | 8.500000 | 4.300000 | 5.760000 |
| 31 | − + + + + | 3.500000 | 2.500000 | 2.500000 | 4.500000 | 8.500000 | 4.300000 | 4.960000 |
| 32 | + + + + + | 4.500000 | 2.500000 | 2.500000 | 4.500000 | 8.500000 | 4.500000 | 4.800000 |
| Degree | Squared bias | Prediction variance | Response noise | Expected error | Average training MSE |
|---|---|---|---|---|---|
| 0 | 1.562500 | 0.050000 | 0.250000 | 1.862500 | 5.000000 |
| 1 | 3.062500 | 0.056250 | 0.250000 | 3.368750 | 2.950000 |
| 2 | 0.000000 | 0.110938 | 0.250000 | 0.360938 | 0.100000 |
| 3 | 0.000000 | 0.154004 | 0.250000 | 0.404004 | 0.050000 |
| 4 | 0.000000 | 0.185150 | 0.250000 | 0.435150 | 0.000000 |
Fit polynomial coefficients to minimize the sum of squared gaps on each dataset’s five targets. The mean function is known only for constructing and evaluating this toy; it does not enter the fitting procedure. For degree 4 five distinct X positions identify a unique interpolant. Curves plot 81 evaluated X positions joined by short straight segments, including both probe positions.
Each training target’s sign is independently equally likely plus or minus. Draw ID minus1 is its five-bit mask, from lowest bit at X = −2 to highest at X = 2. A fresh response at the probe has an independent plus/minus sign and the same amplitude. Its conditional mean is f(X) and noise variance is amplitude². Average over the full32-fit distribution, using variance divisor 32. Expanding squared prediction error under this independence gives squared bias + prediction variance + response noise. Expected error is also computed directly from all 64 combinations.
A single fit’s squared residual is neither squared bias nor repeated-training prediction variance. Training draw does not change the underlying distribution. Noise-free collapses all 32 weighted outcomes to identical targets; degree 2 or higher recovers this quadratic with zero error. At X = 0.5 degree 1 has more squared bias than degree 0, so neither a monotone bias curve nor a universal U-shaped total error is assumed. No sampling estimate, confidence interval, real model-selection procedure, causal conclusion or guarantee about other populations is made.
scikit-learn · Repeated-training bias, variance and noise