AI Grounds

Open AI Grounds on a desktop

These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.

Guided discovery

Inside a Transformer Block

Follow one residual stream through a decoder block.

First residual addition · Selected token 1

The input splits before RMSNorm. One path carries raw x unchanged; the other normalizes it, mixes allowed tokens through attention, then is scaled and added.

Input x(1, 2)Input x, unchangedRMSNormCausalattentionScaled branch(0.200, -0.100)+First sum h(1.200, 1.900)

h = x + attention branch = (1.000000, 2.000000) + (0.200000, -0.100000) = (1.200000, 1.900000)

Then: RMSNorm(h) → MLP → scale → add unchanged h → Output. This disclosed pre-norm block is not a full trained transformer.

Multiply the attention contribution before the first residual addition; input x bypasses it. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Multiply the MLP contribution before the second addition; h bypasses it. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Selected operation: First residual sum

Token 1 · (1.200000, 1.900000).

Feature 11.20Feature 21.90−606

Attention weights for this token: 1.000000, 0.000000. Future token 2 is masked (weight0).

Exact vectors through the complete block

Inputs: token 1 (1.000000, 2.000000); token 2 (-1.000000, 1.000000). Every operation preserves two features. Values use full precision; display rounds to six decimals.

StageToken1Token2
Input x(1.000000, 2.000000)(-1.000000, 1.000000)
RMSNorm1 (also Q,K)(0.632455, 1.264911)(-1.000000, 1.000000)
Projected V(0.200000, -0.100000)(-0.316228, -0.079057)
Scaled attention(0.200000, -0.100000)(-0.174018, -0.084826)
First residual sum h(1.200000, 1.900000)(-1.174018, 0.915174)
RMSNorm2(0.755180, 1.195702)(-1.115368, 0.869455)
MLP ReLU hidden(0.755180, 1.195702)(0.000000, 0.869455)
Scaled MLP(0.270606, 0.283192)(0.086946, 0.260837)
Output = h + MLP branch(1.470606, 2.183192)(-1.087072, 1.176010)
q·k/sqrt(2); causal mask is applied before softmax.
Query tokenKey1 score/weightKey2 score/weight
11.414213 / 1.000000Masked / 0
20.447213 / 0.2754791.414212 / 0.724521
Exact matrices, normalization and sources

RMSNorm(x)=x/sqrt(mean(x²)+0.000001), gain 1.
Wq=Wk=Wo=Wup=identity2×2.
Wv=diag(0.316227829, -0.079056957).
Wdown rows: (0.200000, 0.100000); (-0.100000, 0.300000).

One causal head, two tokens, two features. RMSNorm and the two-layer ReLU MLP operate independently per token. Branch scales are teaching interventions after attention and after the MLP; they are not training. No dropout, learned positional encoding, final normalization, language head or token generation is included. Input hidden states are treated as already positioned. The fixed Wv was chosen so token 1's initial attention branch is(0.2,−0.1), not estimated from training. This pre-norm RMS variant differs from the original post-norm transformer.

Attention Is All You Need · Attention and feed-forward operations
On Layer Normalization · Pre-norm placement
Root Mean Square Layer Normalization