Scaled Dot-Product Attention
P10.attention.02 · Audience: guest, it-ml, language-pro · Prerequisites: What is Attention?
Knowing that each word should borrow meaning from the words that clarify it is one thing; a machine needs an exact, differentiable recipe for deciding how much every word borrows from every other. Scaled dot-product attention is that recipe: each token is given a question to ask, an answer to offer and content to carry, and a short chain of matrix operations turns those three roles into a weighted mixing of the whole sentence. We walk that chain one step at a time, from the first projections to the final context vectors, with every intermediate tensor computed live from the active corpus.
🗣️ From a linguist's perspective: the sentence as a matrix
ⓘ Concept: Q, K, V — three roles for every token
Why it matters — Every contextual embedding a Transformer produces starts from these three projections — asking (Q), advertising (K), and carrying content (V).
Attention input
Reads /api/attention/pipeline — loading
What training does to these maps
Everything above ran on seeded stand-in weights, so you could study the mechanism with nothing hidden. The model store holds the same architecture after training — at its measured budget, for each of the three objectives — and the panel below reads those weights and draws the same sentence's attention beside your stand-in. Nothing here trains: a sentence the stored corpus cannot embed, or a configuration the store does not hold, says so instead of computing.
Trained attention
Reads /api/attention/trained — loading
Stand-in — your controls (d_model 16, d_k 8, seed 1, unmasked)
Computing…
Reading the model store…
The attention demo runs this computation with several heads over a three-layer stack, on the same corpus.
🎓 Interview deep-dive
Causal vs bidirectional masking. Encoders (BERT) attend in both directions; decoders
(GPT) apply a causal mask — token i may only attend to positions ≤ i — implemented by
setting the upper triangle of the score matrix to −∞ before softmax — switch the trained panel
above to causal and its upper triangle is exactly 0. The padding mask in Step 5 is the same
trick applied to a different pattern.
Complexity. Attention is O(n²·d) in sequence length — the reason long-context models need FlashAttention (tiling that never materialises the full n×n matrix in slow memory) and activation checkpointing during training.
📚 Go further
- Vaswani et al. 2017 — Attention Is All You Need
- Jay Alammar — The Illustrated Transformer
- Lilian Weng — Attention? Attention!
- Andrej Karpathy — Let's build GPT
Check your understanding
Self-check
The raw score between token i and token j is computed how?
Self-check
Why divide the scores by √d_k before softmax?
Self-check
What does each row of the weight matrix sum to, and why does that matter?
Open in Labs: The attention matrix
Ask the mentor about this module
Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.
Keeping your files on this device
Off by default. The mentor always gets your file; this only decides whether your own copy stays here. Copies live in this browser only - they do not follow you to another device, and clearing site data removes them.
🎓 Practice ladder
1 graded rung · ~12 minWalk the five steps yourself. The rung is a three-panel workspace: instructions, a code editor, and output + test results. Run checks the visible tests; Submit grades against hidden edge cases — the weights must be row-stochastic, the scores scaled by √dₖ before softmax, and the context a weighted sum of the value rows.
Rung 1 — One attention head, five steps
Loading exercise…
Capstone — Explain attention
Prompt. Explain the attention mechanism to someone who has never studied machine learning. Your explanation should cover: (1) what attention scores measure, (2) why we divide by √dₖ, and (3) how softmax turns scores into weights. Write your own answer first, then reveal the model answer.
💡 Model answer
What attention scores measure. Every word sends out a question (its Query vector) asking the rest of the sentence for relevant context, and simultaneously publishes an advert (its Key vector) describing what it has to offer. An attention score is the dot product of one word's question with another word's advert — a high score means the two are a good match, so the first word should attend to the second.
Why divide by √dₖ. With many dimensions (large dₖ) each dot product sums many products, so its magnitude grows with dₖ. Very large scores push softmax into saturation — one weight becomes ≈ 1 and the rest ≈ 0 — which slows learning. Dividing by √dₖ scales the scores back to a manageable range and keeps the distribution spread across positions.
How softmax turns scores into weights. Softmax exponentiates every score (making them positive) and divides each by the total, so they sum to exactly 1 — a probability distribution. Each token's context vector is then a weighted average of all Value vectors: a token that assigns 0.7 to the subject takes 70 % of its new representation from the subject's value vector.
Try it yourself
A scratch console for this page's ideas — ungraded, nothing you run here is recorded.
Scratch console
A scratch console with the scientific stack (pandas, numpy, scikit-learn). Runs on the server — no network, resource-limited and measured.
Output appears here.
My notes on this module
Loading your notes...
Where next?
Later in Attention
Go up a level