AI Grounds

Open AI Grounds on a desktop

These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.

Guided discovery

Regularization Lab

Choose a penalty. Then fit the weights.

Your dataset

Eight authored cases have two features and fixed targets +1 or −1. Score = w1×x1 + w2×x2; bias stays zero. This toy scorer uses squared error and then a sign cutoff, not logistic probabilities. Predict +1 when score≥0, including exact ties, and −1 otherwise. Cases 6 and 8 share coordinates but have opposite targets; both cannot be correct under one deterministic score.

-1.5-1.5001.51.51,3: +1,+1 at (1,0); fixed targets, two cases1,3: +1,+12,4: −1,−1 at (-1,0); fixed targets, two cases2,4: −1,−15,7: +1,+1 at (0,1); fixed targets, two cases5,7: +1,+16,8: −1,+1 at (0,-1); fixed targets, two cases6,8: −1,+1x1x2

Score-zero boundary: 1×x1 + 0.5×x2 = 0. Boundary scores tie and predict +1.

Circles show single-target groups, a square mixed targets. Text lists exact case IDs and fixed labels; marker color does not represent model confidence. Teal region predicts +1, orange −1; boundary ties predict +1.

Signed coefficient for feature x1. Manual edits immediately change scores, data loss and the current boundary. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Signed coefficient for feature x2. Manual edits immediately change scores, data loss and the current boundary. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Mode changes keep current weights fixed.

λ multiplies the chosen penalty; None has zero penalty at every strength. A strength edit keeps current weights fixed. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Fit checks all 81×81=6,561 weight pairs in −2..2, step 0.05. It minimizes data loss plus the current penalty on this grid; it does not prove the unrestricted continuous optimum. Among ties it applies the lowest Weight 1, then Weight 2.

Unregularized · Weights (1, 0.5) · None · Strength 0.25 · 7/8 correct, 87.5%

Current weights are a best grid fit. 1 best pair: (1, 0.5). Minimum objective 0.1875. Fit applies (1, 0.5).

Actual weight, mode or strength changes clear stale answers; unchanged controls and a fit that leaves weights unchanged preserve them. All accuracy here describes the same eight training cases; no held-out evidence is supplied.

Fixed-case evidence

Both cases at every site remain separate. A zero score predicts +1. Each row contributes ½×(score−target)²; data loss averages all eight terms.
IDx1x2TargetScorePredictedTie?Correct?½ residual²
110111NoYes0
2-10-1-1-1NoYes0
310111NoYes0
4-10-1-1-1NoYes0
50110.51NoYes0.125
60-1-1-0.5-1NoYes0.125
70110.51NoYes0.125
80-11-0.5-1NoNo1.125

Data loss = (0 + 0 + 0 + 0 + 0.125 + 0.125 + 0.125 + 1.125) / 8 = 0.1875. Penalty = 0 (no penalty) = 0. Objective = 0.1875 + 0 = 0.1875.

Data loss

Mean half-squared error over all eight fixed cases.

0.1875

Σ ½(score − target)² / 8 = SSE / 16

Includes every case, whether its sign prediction is correct or not.

Penalty

The chosen cost of coefficient magnitudes.

0

None: 0

Identical λ does not imply equal L1 and L2 penalty magnitudes.

Objective

Data loss plus the current penalty.

0.1875

0.1875 + 0

Fit minimizes this quantity on the declared weight grid, not accuracy.

Construction and limits

Targets are five +1 and three −1 labels. Cases 6 and 8 have identical (0,−1) features and opposite targets. Inputs are fixed, same-scale toy columns; there is no intercept fit. We fit signed scores by squared error and then use a sign cutoff for an illustrative decision boundary. This is not logistic regression, calibrated probability or a recommended application model.

Exact grid parameters use twentieths: weights −40..40 divided by 20, strength 0..20 divided by 20. Objective arithmetic has an exact integer numerator over 16,000, so tied grid minima are identified without a numerical tolerance. Fit searches every pair, reports all best pairs and applies their ascending Weight 1/Weight 2 order. At L2 strength 0.30 the tied pairs are (0.60,0.30) and (0.65,0.30), both objective 0.305; Fit chooses (0.60,0.30).

L1 uses λ times absolute weights; L2 uses λ/2 times squared weights. Data loss is SSE/(2n), n=8. Other libraries or texts can use different scale factors, so λ must be interpreted with its declared formula. L1 can create zeros; L2 often shrinks weights, but neither guarantees improved accuracy or useful feature selection. Grid approximation, feature scaling and different datasets can change coefficients or boundary ratios. The experiment’s exact common half-scale is one specific L2 case.

Penalty edits leave current weights fixed; only a weight edit or Fit changes them. Training data loss and training accuracy are different from objective and from held-out generalization. No validation data, optimizer path, learned feature encoder, automatic λ selection or bias penalty is supplied. An all-zero scorer has no unique separating line, but its explicit +1 tie rule still defines predictions and accuracy. Reset/prediction restores Unregularized for the current experiment; free-exploration Reset begins Experiment 1.

scikit-learn · Squared-error, L1 and L2 scaling convention
Stanford STATS 202 · Shrinkage and sparsity