AI Grounds
Open AI Grounds on a desktop
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
AI Grounds
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
Guided discovery
Rearrange tokens. Compare identity with position signals.
Three authored identities have fixed two-number vectors: A=(1,0), B=(0,1), C=(−1,0). Positions start at zero. Query token selects an identity and follows it when reordered. These vectors do not represent learned word meanings. No mask, attention weights or model outputs run here.
Purple bar / zero dot: current score. Gray dashed tick: A B C at start 0 reference, using the same Position signal and Query token. Signed raw scores are unscaled dot products: multiply matching Q/K coordinates, then add. They are not probabilities or attention weights.
Shift every zero-based position by the same amount; token order and identity stay fixed. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.
No signal: Q and K use unchanged token vectors.
A radian is an angle unit. Sine and cosine provide the shown position coordinates. These are two-dimensional, one-frequency examples; learned query/key projections are replaced with fixed identity projections. All actual control/scenario changes clear stale explanations and transfer answers; reapplying the same state preserves them.
Selected token’s actual zero-based position.
0
start position + slot index
Query A has current Q vector (1, 0).
How many per-key scores differ from reference.
0
compare same key identity, mode and query
Reference fixes A B C at start 0. A 10⁻¹² tolerance handles floating arithmetic.
Maximum absolute raw-score difference.
0
max |current score − reference score|
Changing both order and start compares both factors. Display rounding is not calculation.
Rows stay keyed by token identity so a reorder is easy to compare. Absolute adds the position’s sine/cosine pair; Rotary rotates Q/K instead of adding a vector. Each raw score uses the selected query’s Q and that row’s K. Fixed identity projections make Q and K equal for the same token here; real learned projections need not.
| Key token | Position | Base vector | Q / K vector | Vector length | Raw score | Reference score | Difference |
|---|---|---|---|---|---|---|---|
| A | 0 | (1, 0) | (1, 0) | 1 | 1 | 1 | 0 |
| B | 1 | (0, 1) | (0, 1) | 1 | 0 | 0 | 0 |
| C | 2 | (-1, 0) | (-1, 0) | 1 | -1 | -1 | 0 |
All six permutations contain the same three token identities once. Start position 0–4 gives three slots at start, start+1 and start+2. The query follows its selected identity. No signal leaves the base vectors unchanged. Absolute adds (sin(p),cos(p)) in this single sine/cosine pair, then applies identity query/key projections. Rotary applies R(p) to the projected Q and K; this is a 2D RoPE example at one radian per position, not an addition to the token embedding. Full models use many dimensions, frequencies and learned projections.
For Rotary, Q at position m dotted with K at n equals baseQ dotted with R(n−m)baseK. A common shift preserves n−m and therefore the score for fixed content/order. Rotation preserves vector length; all raw unit vectors here have rotary length 1 and self-score 1. Absolute addition need not preserve length or scores after a shift. Rotary scores combine content and offsets and can oscillate; this demo makes no monotonic distance-decay claim.
No signal gives permutation-invariant per-token pair scores in this isolated unmasked, content-only calculation. Position-indexed vectors still move with the permutation. Masks and other position signals can convey order; do not generalize this result to every model without this particular encoding. The reference fixes both order A B C and start 0, so changing both controls compares both factors. Presets restore all controls to A B C, start 0, query A in the selected rule. Reset/predictions restore No position for the current experiment; free Reset starts Experiment 1.
Raw dot scores are unscaled, can be negative or exceed 1, and are bounded by −4..4 here because the largest encoded length is 2. No softmax, probabilities, attention-weighted values, model output, tokenizer, training, semantic answer or context-length extrapolation is supplied. Position signals make order information available; they do not prove learned language understanding. Values smaller than 10⁻¹² display as zero; Changed scores uses that tolerance, not six-decimal rounded equality.
Attention Is All You Need · Absolute sinusoidal encoding