AI Grounds

Open AI Grounds on a desktop

These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.

Guided discovery

Train/Test Split & Generalization Lab

Fit on one group; evaluate on another.

Your dataset

Twelve fixed toy pairs, split into disjoint groups of four. Train fits the candidates; Validation guides choice; Test evaluates afterward. Split changes membership, not pair values.

Fitting rows 4 · Train MSE 0.000 · Validation MSE 8.500 · Test not evaluated. Nearest neighbor copies the nearest fitting target.

Model

Fitting rows

-2410162228036911YX

Circle = Train (4); diamond = Validation (4); triangle = Test (4), withheld. Hollow square = candidate prediction at each of the twelve X positions. Line draws a straight prediction rule; Nearest neighbor’s predictions are unconnected. Coincident square/observation shapes remain visible.

Use Validation for model choice, then evaluate Test. Changing split, model or fitting source hides the Test targets and clears completion. This repeatable toy cannot recreate untouched real test data by resetting or hiding previously seen targets.

MSE = sum of squared observed-minus-predicted gaps / 4 in each original group. Fitting uses 4 rows; group sizes remain 4, 4 and 4.

Train MSE

Average squared error on the four original Train rows.

0.000

Train squared-gap sum / 4

Validation MSE

Compare candidates using the four Validation rows.

8.500

Validation squared-gap sum / 4

Unused by the fit; used for model choice.

Test MSE

Evaluate the chosen procedure on the four Test rows.

Withheld

Choose Evaluate test after model choice.

Keep these targets out of fitting and choices.

These are fixed constructed examples, not random population samples. A finite held-out error can be higher, lower or equal to another group’s error; it does not guarantee future accuracy or establish causation.

Twelve pairs and assignments
The pairs stay fixed; Test target values and errors are withheld before evaluation.
PointRoleXObserved YUsed to fitPredicted YResidualSquared
1Train01.000Yes1.0000.0000.000
2Validation15.000No1.000+4.00016.000
3Test2WithheldNo8.000WithheldWithheld
4Train38.000Yes8.0000.0000.000
5Validation47.000No8.000−1.0001.000
6Test5WithheldNo15.000WithheldWithheld
7Train615.000Yes15.0000.0000.000
8Validation714.000No15.000−1.0001.000
9Test8WithheldNo17.000WithheldWithheld
10Train917.000Yes17.0000.0000.000
11Validation1021.000No17.000+4.00016.000
12Test11WithheldNo17.000WithheldWithheld
Construction and sources

The twelve paired observations are fixed. Line recomputes the least-squares slope and intercept from the selected fitting rows. Nearest neighbor copies the Y of the closest fitting X, with a lower Point ID winning an equal-distance tie. Both are actual fitting/prediction rules, not scripted score changes.

Train only uses four Train rows. Validation scores can guide candidate choice; Test is evaluated later without refitting. All rows (leak) deliberately uses all twelve targets, including the groups named Validation and Test. Their group labels stay visible for comparison, but their evaluation is contaminated. Each group’s MSE still averages its own four errors.

Changing a setting hides Test evaluation and changes the fit when relevant, but cannot make previously observed Test targets new. Repeatedly consulting the same Test results to choose settings gives them a validation role; real final evaluation then needs untouched data. Learned preprocessing also belongs inside the training-only fitting procedure. This toy has no preprocessing operation, cross-validation, automatic model search or confidence interval.

Left training evaluates higher X values outside its fitting range. Appropriate real splits depend on the task, dependencies, groups and time; one fixed partition or low score does not establish representative generalization. Full-precision errors precede three-decimal display. MSE uses squared Y units, not classification accuracy or a causal quantity.

scikit-learn · Fitting separation and data leakage
scikit-learn · Validation-based choice and later Test evaluation