AI Grounds
Open AI Grounds on a desktop
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
AI Grounds
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
Guided discovery
Rescale a vector. Compare RMS scaling with mean centering.
One authored activation vector has three feature coordinates. Both rules use the same inputs, Epsilon and shared Gain γ; no batch statistics, running averages or training are involved. RMS means root mean square: square all inputs, average the three squares, then take the square root. RMSNorm divides the original coordinates; LayerNorm first subtracts their mean.
Bars start at zero on the shared signed output scale. Dots mark a defined zero output. Undefined means a zero denominator; it has no bar or invented zero marker. Exact values and Positive-reference outputs follow in the construction table.
Feature coordinate 1. Both rules and every statistic update after an edit. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.
Feature coordinate 2. Both rules and every statistic update after an edit. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.
Feature coordinate 3. Both rules and every statistic update after an edit. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.
One nonnegative shared scalar multiplies all normalized outputs afterward. Real models can learn separate gains per feature. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.
RMSNorm = gain × x / sqrt(mean(x²) + epsilon)
LayerNorm = gain × (x − mean) / sqrt(variance + epsilon)
mean = sum(x)/3; variance = sum((x − mean)²)/3; bias β=0
Input mean 2 · RMS denominator 2.160247 · LayerNorm denominator 0.816497. LayerNorm output mean: 0.
Epsilon is inside the square root; gain is applied after normalization. A zero denominator is undefined even when gain is zero. Actual input, gain, epsilon or preset changes clear stale explanations and transfer answers; reapplying the same state preserves them. The reference always uses Positive inputs 1,2,3 with the current gain and epsilon.
Root mean square of the original input.
2.160247
sqrt(mean(x²))
Mean square=4.666667. This statistic is defined at zero even if normalization is not.
Average of the three RMSNorm outputs.
0.92582
gain × input mean / RMS denominator
RMSNorm does not subtract the mean. Shared-gain LayerNorm with zero bias has mean zero when defined.
Root mean square after RMSNorm and gain.
1
gain × sqrt(meanSquare / (meanSquare + epsilon))
No general unit-RMS guarantee. Small positive epsilon can round a slightly smaller value to 1.
Population variance averages all three squared centered deviations; it does not use a sample-size correction. Mean square averages squares without centering. The table shows both constructions and compares outputs with Positive 1,2,3 under the current gain and epsilon. Display rounds to six decimals; calculations use unrounded values.
| Coordinate | Input | Square | Centered | RMSNorm output | LayerNorm output | RMS reference | Layer reference |
|---|---|---|---|---|---|---|---|
| x1 | 1 | 1 | -1 | 0.46291 | -1.224744 | 0.46291 | -1.224744 |
| x2 | 2 | 4 | 0 | 0.92582 | 0 | 0.92582 | 0 |
| x3 | 3 | 9 | 1 | 1.38873 | 1.224744 | 1.38873 | 1.224744 |
Inputs are three integer feature values −6..6. They are normalized together as one vector, independently of other tokens or samples. Gain 0..2 in steps of 0.5 is one nonnegative scalar shared across coordinates and both rules; real elementwise gains may differ or be negative and can be learned. This toy has no gain training or additive bias. Epsilon choices 0,0.000001 and 1 are explicit demo settings, not a library precision-dependent default. Zero exposes a mathematical boundary; 1 is a large contrast, not a recommended training value.
RMSNorm uses sqrt(mean(x²)+epsilon), not sqrt(mean(x²))+epsilon. LayerNorm uses sqrt(mean((x−mean)²)+epsilon) and the original minus mean as numerator. Both apply gain afterward. When a denominator is zero, normalization and its gained output are Undefined, even with gain zero. All-zero input has zero input RMS but undefined normalization at epsilon zero. With positive epsilon, both zero-vector outputs are zero. A constant nonzero vector at epsilon zero has defined RMSNorm but undefined LayerNorm; positive epsilon makes constant-vector LayerNorm output zero.
For a nonzero vector and epsilon zero, multiplying all inputs by a positive factor cancels exactly in RMS normalization. A fixed positive epsilon generally breaks exact rescaling invariance. A negative input factor flips the normalized sign at epsilon zero, so do not call that identical output. A common additive shift preserves LayerNorm’s centered deviations and variance when its denominator is nonzero; it need not preserve RMSNorm. If the input already has mean zero, both rules agree with the same gain and epsilon, without RMSNorm performing centering.
Normalized RMS is sqrt(meanSquare/(meanSquare+epsilon)) when defined; output RMS is the shared nonnegative gain times that value. Epsilon zero gives normalized RMS 1 for nonzero inputs; positive epsilon lowers it. Neither gain nor epsilon guarantees a final RMS of 1. This lesson demonstrates arithmetic, not measured training speed, transformer residual-stream behavior, long-context performance or guaranteed stable training. Presets restore gain 1 and epsilon 0.000001. Reset/predictions restore Positive for the current experiment; free Reset starts Experiment 1.
PyTorch · RMSNorm denominator and elementwise gain