Multi-Head Attention
P10.attention.03 · Audience: guest, it-ml, language-pro · Prerequisites: Scaled Dot-Product Attention
A single attention pass produces only one weight pattern per sentence, so it has to squeeze grammar, meaning and position into that one pattern at once. A careful reader tracks these threads separately, and the Transformer gains the same faculty by running several smaller attention heads in parallel, each free to specialise in its own kind of relationship. Here we see how the heads divide the projection weights without ever splitting the input, and how a final learned projection recombines their separate readings into one representation.
🗣️ From a linguist's perspective: simultaneous levels of analysis
ⓘ Concept: Why multiple attention heads?
Why it matters — One head averages all relationship types into one pattern; separate heads are free to specialise — BERT-base uses 12, GPT-3 uses 96.
Reference: Vaswani et al. 2017, §3.2.2
Computing…
For several heads over a three-layer stack on your corpus, open the attention demo.
📚 Go further
- Vaswani et al. 2017, §3.2.2 — Attention Is All You Need
- Jay Alammar — The Illustrated Transformer
- Harvard NLP — The Annotated Transformer
- Lilian Weng — Attention? Attention!
Check your understanding
Self-check
In multi-head attention, is the input matrix X split across heads?
Self-check
If d_model = 16 and num_heads = 4, what is d_k?
Self-check
What would be lost without the final W_O projection?
Ask the mentor about this module
Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.
Keeping your files on this device
Off by default. The mentor always gets your file; this only decides whether your own copy stays here. Copies live in this browser only - they do not follow you to another device, and clearing site data removes them.
Try it yourself
A scratch console for this page's ideas — ungraded, nothing you run here is recorded.
Scratch console
A scratch console with the scientific stack (pandas, numpy, scikit-learn). Runs on the server — no network, resource-limited and measured.
Output appears here.
My notes on this module
Loading your notes...
Where next?
Later in Attention
This module unlocks
Go up a level