AI Grounds
Open AI Grounds on a desktop
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
AI Grounds
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
Guided discovery
Follow one residual stream through a decoder block.
The input splits before RMSNorm. One path carries raw x unchanged; the other normalizes it, mixes allowed tokens through attention, then is scaled and added.
h = x + attention branch = (1.000000, 2.000000) + (0.200000, -0.100000) = (1.200000, 1.900000)
Then: RMSNorm(h) → MLP → scale → add unchanged h → Output. This disclosed pre-norm block is not a full trained transformer.
Multiply the attention contribution before the first residual addition; input x bypasses it. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.
Multiply the MLP contribution before the second addition; h bypasses it. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.
Token 1 · (1.200000, 1.900000).
Attention weights for this token: 1.000000, 0.000000. Future token 2 is masked (weight0).
Inputs: token 1 (1.000000, 2.000000); token 2 (-1.000000, 1.000000). Every operation preserves two features. Values use full precision; display rounds to six decimals.
| Stage | Token1 | Token2 |
|---|---|---|
| Input x | (1.000000, 2.000000) | (-1.000000, 1.000000) |
| RMSNorm1 (also Q,K) | (0.632455, 1.264911) | (-1.000000, 1.000000) |
| Projected V | (0.200000, -0.100000) | (-0.316228, -0.079057) |
| Scaled attention | (0.200000, -0.100000) | (-0.174018, -0.084826) |
| First residual sum h | (1.200000, 1.900000) | (-1.174018, 0.915174) |
| RMSNorm2 | (0.755180, 1.195702) | (-1.115368, 0.869455) |
| MLP ReLU hidden | (0.755180, 1.195702) | (0.000000, 0.869455) |
| Scaled MLP | (0.270606, 0.283192) | (0.086946, 0.260837) |
| Output = h + MLP branch | (1.470606, 2.183192) | (-1.087072, 1.176010) |
| Query token | Key1 score/weight | Key2 score/weight |
|---|---|---|
| 1 | 1.414213 / 1.000000 | Masked / 0 |
| 2 | 0.447213 / 0.275479 | 1.414212 / 0.724521 |
RMSNorm(x)=x/sqrt(mean(x²)+0.000001), gain 1.
Wq=Wk=Wo=Wup=identity2×2.
Wv=diag(0.316227829, -0.079056957).
Wdown rows: (0.200000, 0.100000); (-0.100000, 0.300000).
One causal head, two tokens, two features. RMSNorm and the two-layer ReLU MLP operate independently per token. Branch scales are teaching interventions after attention and after the MLP; they are not training. No dropout, learned positional encoding, final normalization, language head or token generation is included. Input hidden states are treated as already positioned. The fixed Wv was chosen so token 1's initial attention branch is(0.2,−0.1), not estimated from training. This pre-norm RMS variant differs from the original post-norm transformer.
Attention Is All You Need · Attention and feed-forward operations