Positional Encoding
P10.positional-encoding.01 · Audience: guest, it-ml, language-pro · Prerequisites: Embeddings
The embeddings we have built so far tell the model what each word means, but nothing about where it stands — and attention on its own cannot tell "the dog bit the man" from "the man bit the dog", because shuffling the words leaves every dot product unchanged. Yet word order is precisely what carries who did what to whom. So before attention can read a sentence rather than a bag of words, every embedding needs a second ingredient: a fingerprint of its position. In this module we build that fingerprint — the sinusoidal positional encoding — and add it to the token embeddings of the active corpus, so that the same word at different positions finally gets a different vector.
🗣️ From a linguist's perspective: order as signal
ⓘ Concept: The permutation-invariance problem
Why it matters — Attention is a content-only weighted sum: shuffle the input rows and the dot products are unchanged. Without position information a Transformer reads a bag of words.
Reference: Vaswani et al. 2017, §3.5
Computing…
🎓 Interview deep-dive
Permutation equivariance, precisely. Without PE, attention satisfies Attention(P·X) = P·Attention(X) for any permutation matrix P: the output is a content-only weighted sum over an unordered set of (key, value) pairs, so reordering inputs merely reorders outputs. PE breaks that symmetry on purpose — it is the only thing telling the model that position 3 differs from position 7.
📚 Go further
- Vaswani et al. 2017, §3.5 — Attention Is All You Need
- Amirhossein Kazemnejad — Transformer Architecture: The Positional Encoding
- Harvard NLP — The Annotated Transformer
- Andrej Karpathy — Let's build GPT
Check your understanding
Self-check
Without position information, can a Transformer distinguish the same words in a different order?
Self-check
For d_model = 4, what is PE(1, 0) to four decimals?
Self-check
What is PE(0, 2i) for every i?
Self-check
A 4-token sentence with d_model = 8: what is the shape after adding PE?
Self-check
Why do identical words at different positions end up with different vectors?
Self-check
For d_model = 4, what is PE(0, 1)?
Ask the mentor about this module
Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.
Keeping your files on this device
Off by default. The mentor always gets your file; this only decides whether your own copy stays here. Copies live in this browser only - they do not follow you to another device, and clearing site data removes them.
🎓 Practice ladder
1 graded rung · ~6 minBuild the encoding yourself. The rung is a three-panel workspace: instructions, a code editor, and output + test results. Run checks the visible tests; Submit grades against hidden edge cases — even dimensions use sine, odd use cosine, the sine/cosine pair shares a frequency, and position 0 is all sines-to-0 / cosines-to-1.
Rung 1 — Fill the 4×4 PE table
Loading exercise…
Try it yourself
A scratch console for this page's ideas — ungraded, nothing you run here is recorded.
Scratch console
A scratch console with the scientific stack (pandas, numpy, scikit-learn). Runs on the server — no network, resource-limited and measured.
Output appears here.
My notes on this module
Loading your notes...
Where next?
Later in Positional Encoding
This module unlocks
Go up a level