How Positional Encoding Works
Position Matters
Why word order changes meaning
Sentence
Dog bites man
Meaning
Normal
Attn change
Same values!
"Dog bites man" — the normal sentence. The attention matrix shows dot products between all word pairs. Try the other orderings and watch what happens.
Step 1 of 6: Position Matters
Why word order changes meaning
Sentence
Dog bites man
Meaning
Normal
Attn change
Same values!
"Dog bites man" — the normal sentence. The attention matrix shows dot products between all word pairs. Try the other orderings and watch what happens.
Step 2 of 6: Sinusoidal Encoding
Sin and cos waves encode position
Position
10
Dim
4 (sin)
Angle
1.58
PE value
1.000
Low dimensions oscillate rapidly — notice the tight vertical stripes on the left side of the heatmap. These encode fine-grained position differences.
Step 3 of 6: Frequency Bands
Different dimensions, different scales
Dimension
0
Frequency
1.0000
Wavelength
6.3
Low dimensions have high frequency — they change rapidly across positions, encoding fine-grained position differences.
Step 4 of 6: Relative Distance
Why sinusoidal encodings work
pos_i
3
pos_j
8
Dot product
6.14
Distance
5
As distance grows, the dot product decreases smoothly. The sinusoidal design ensures similarity depends only on relative distance.
Step 5 of 6: RoPE
Rotary position embedding via 2D rotation
q pos
3
k pos
7
q·k
1.18
|q−k|
4
RoPE rotates q and k by different amounts (Δθ depends on |3−7|=4). The dot product depends only on this relative distance — not on absolute position. Move both sliders by the same offset and watch q·k stay constant!
Step 6 of 6: Sinusoidal vs RoPE
Comparing the two approaches
|m−n|
6
Sin dot
3.78
RoPE dot
1.97
At large distances, both show low similarity. RoPE's advantage: since it encodes relative position directly into Q·K, it generalizes better to sequence lengths not seen during training. That's why LLaMA, Mistral, and Gemma all use RoPE.