AI Grounds
Open AI Grounds on a desktop
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
AI Grounds
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
Guided discovery
Fit on one group; evaluate on another.
Twelve fixed toy pairs, split into disjoint groups of four. Train fits the candidates; Validation guides choice; Test evaluates afterward. Split changes membership, not pair values.
Fitting rows 4 · Train MSE 0.000 · Validation MSE 8.500 · Test not evaluated. Nearest neighbor copies the nearest fitting target.
Circle = Train (4); diamond = Validation (4); triangle = Test (4), withheld. Hollow square = candidate prediction at each of the twelve X positions. Line draws a straight prediction rule; Nearest neighbor’s predictions are unconnected. Coincident square/observation shapes remain visible.
Use Validation for model choice, then evaluate Test. Changing split, model or fitting source hides the Test targets and clears completion. This repeatable toy cannot recreate untouched real test data by resetting or hiding previously seen targets.
MSE = sum of squared observed-minus-predicted gaps / 4 in each original group. Fitting uses 4 rows; group sizes remain 4, 4 and 4.
Average squared error on the four original Train rows.
0.000
Train squared-gap sum / 4
Compare candidates using the four Validation rows.
8.500
Validation squared-gap sum / 4
Unused by the fit; used for model choice.
Evaluate the chosen procedure on the four Test rows.
Withheld
Choose Evaluate test after model choice.
Keep these targets out of fitting and choices.
These are fixed constructed examples, not random population samples. A finite held-out error can be higher, lower or equal to another group’s error; it does not guarantee future accuracy or establish causation.
| Point | Role | X | Observed Y | Used to fit | Predicted Y | Residual | Squared |
|---|---|---|---|---|---|---|---|
| 1 | Train | 0 | 1.000 | Yes | 1.000 | 0.000 | 0.000 |
| 2 | Validation | 1 | 5.000 | No | 1.000 | +4.000 | 16.000 |
| 3 | Test | 2 | Withheld | No | 8.000 | Withheld | Withheld |
| 4 | Train | 3 | 8.000 | Yes | 8.000 | 0.000 | 0.000 |
| 5 | Validation | 4 | 7.000 | No | 8.000 | −1.000 | 1.000 |
| 6 | Test | 5 | Withheld | No | 15.000 | Withheld | Withheld |
| 7 | Train | 6 | 15.000 | Yes | 15.000 | 0.000 | 0.000 |
| 8 | Validation | 7 | 14.000 | No | 15.000 | −1.000 | 1.000 |
| 9 | Test | 8 | Withheld | No | 17.000 | Withheld | Withheld |
| 10 | Train | 9 | 17.000 | Yes | 17.000 | 0.000 | 0.000 |
| 11 | Validation | 10 | 21.000 | No | 17.000 | +4.000 | 16.000 |
| 12 | Test | 11 | Withheld | No | 17.000 | Withheld | Withheld |
The twelve paired observations are fixed. Line recomputes the least-squares slope and intercept from the selected fitting rows. Nearest neighbor copies the Y of the closest fitting X, with a lower Point ID winning an equal-distance tie. Both are actual fitting/prediction rules, not scripted score changes.
Train only uses four Train rows. Validation scores can guide candidate choice; Test is evaluated later without refitting. All rows (leak) deliberately uses all twelve targets, including the groups named Validation and Test. Their group labels stay visible for comparison, but their evaluation is contaminated. Each group’s MSE still averages its own four errors.
Changing a setting hides Test evaluation and changes the fit when relevant, but cannot make previously observed Test targets new. Repeatedly consulting the same Test results to choose settings gives them a validation role; real final evaluation then needs untouched data. Learned preprocessing also belongs inside the training-only fitting procedure. This toy has no preprocessing operation, cross-validation, automatic model search or confidence interval.
Left training evaluates higher X values outside its fitting range. Appropriate real splits depend on the task, dependencies, groups and time; one fixed partition or low score does not establish representative generalization. Full-precision errors precede three-decimal display. MSE uses squared Y units, not classification accuracy or a causal quantity.
scikit-learn · Fitting separation and data leakage