Guided pathThis is part of Understand Transformers & BERTBack to the path

Positional Encoding

P10.positional-encoding.01 · Audience: guest, it-ml, language-pro · Prerequisites: Embeddings

Customise the corpus these labs use
Real LLM grading for this pageLLM grading (this page):

The embeddings we have built so far tell the model what each word means, but nothing about where it stands — and attention on its own cannot tell "the dog bit the man" from "the man bit the dog", because shuffling the words leaves every dot product unchanged. Yet word order is precisely what carries who did what to whom. So before attention can read a sentence rather than a bag of words, every embedding needs a second ingredient: a fingerprint of its position. In this module we build that fingerprint — the sinusoidal positional encoding — and add it to the token embeddings of the active corpus, so that the same word at different positions finally gets a different vector.

Step 1 / 5 — The Permutation Problem
🗣️ From a linguist's perspective: order as signal
“The dog bit the man” and “the man bit the dog” contain the same words — only their order carries who did what to whom. Any representation that ignores order throws away syntax itself.
ⓘ Concept: The permutation-invariance problem
The Block-4 embeddings encode what each word means, not where it stands — a repeated word gets literally identical rows. Pick a sentence with a repeated word and see for yourself.

Why it matters — Attention is a content-only weighted sum: shuffle the input rows and the dot products are unchanged. Without position information a Transformer reads a bag of words.

Reference: Vaswani et al. 2017, §3.5

Computing…

🎓 Interview deep-dive

Permutation equivariance, precisely. Without PE, attention satisfies Attention(P·X) = P·Attention(X) for any permutation matrix P: the output is a content-only weighted sum over an unordered set of (key, value) pairs, so reordering inputs merely reorders outputs. PE breaks that symmetry on purpose — it is the only thing telling the model that position 3 differs from position 7.

📚 Go further

Check your understanding

Self-check

Without position information, can a Transformer distinguish the same words in a different order?

Self-check

For d_model = 4, what is PE(1, 0) to four decimals?

Self-check

What is PE(0, 2i) for every i?

Self-check

A 4-token sentence with d_model = 8: what is the shape after adding PE?

Self-check

Why do identical words at different positions end up with different vectors?

Self-check

For d_model = 4, what is PE(0, 1)?

Ask the mentor about this module

Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.

Images, PDF or text. Kept on this device only.
Keeping your files on this device

Off by default. The mentor always gets your file; this only decides whether your own copy stays here. Copies live in this browser only - they do not follow you to another device, and clearing site data removes them.

Ctrl/Cmd + Enter to send

🎓 Practice ladder

1 graded rung · ~6 min

Build the encoding yourself. The rung is a three-panel workspace: instructions, a code editor, and output + test results. Run checks the visible tests; Submit grades against hidden edge cases — even dimensions use sine, odd use cosine, the sine/cosine pair shares a frequency, and position 0 is all sines-to-0 / cosines-to-1.

Rung 1 — Fill the 4×4 PE table

Loading exercise…

Try it yourself

A scratch console for this page's ideas — ungraded, nothing you run here is recorded.

Scratch console

A scratch console with the scientific stack (pandas, numpy, scikit-learn). Runs on the server — no network, resource-limited and measured.

Output appears here.

My notes on this module

Loading your notes...

Where next?

Continue

Modern Positional Encodings →

Positional Encoding · P10

Later in Positional Encoding

This module unlocks

Positional Encoding — TransformerLab