Guided pathThis is part of Understand Transformers & BERTBack to the path

Multi-Head Attention

P10.attention.03 · Audience: guest, it-ml, language-pro · Prerequisites: Scaled Dot-Product Attention

Customise the corpus these labs use

A single attention pass produces only one weight pattern per sentence, so it has to squeeze grammar, meaning and position into that one pattern at once. A careful reader tracks these threads separately, and the Transformer gains the same faculty by running several smaller attention heads in parallel, each free to specialise in its own kind of relationship. Here we see how the heads divide the projection weights without ever splitting the input, and how a final learned projection recombines their separate readings into one representation.

Step 1 / 6 — Motivation: Why Multiple Heads?
🗣️ From a linguist's perspective: simultaneous levels of analysis
A linguist reads one sentence at several levels at once — phonology, syntax, semantics, pragmatics. Each analysis attends to different words for different reasons. Multi-head attention gives the model that faculty: each head is one analytical reading, run in parallel.
ⓘ Concept: Why multiple attention heads?
A single head must compress every kind of relationship — syntactic, semantic, positional — into one weight pattern. Multiple heads, each in a smaller d_k-dimensional subspace, are each free to specialise. Below the controls, the trained model's own heads read the selected sentence.
MultiHead(X)=Concat(head1,…,headH) WO,headh=Attention(QWQh,KWKh,VWVh)\text{MultiHead}(X) = \text{Concat}(\text{head}_1,\dots,\text{head}_H)\,W_O,\qquad \text{head}_h = \text{Attention}(QW_{Q_h}, KW_{K_h}, VW_{V_h})

Why it matters — One head averages all relationship types into one pattern; separate heads are free to specialise — BERT-base uses 12, GPT-3 uses 96.

Reference: Vaswani et al. 2017, §3.2.2

Computing…

Open in Labs: Heads split and mixed

For several heads over a three-layer stack on your corpus, open the attention demo.

📚 Go further

Check your understanding

Self-check

In multi-head attention, is the input matrix X split across heads?

Self-check

If d_model = 16 and num_heads = 4, what is d_k?

Self-check

What would be lost without the final W_O projection?

Ask the mentor about this module

Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.

Images, PDF or text. Kept on this device only.
Keeping your files on this device

Off by default. The mentor always gets your file; this only decides whether your own copy stays here. Copies live in this browser only - they do not follow you to another device, and clearing site data removes them.

Ctrl/Cmd + Enter to send

Try it yourself

A scratch console for this page's ideas — ungraded, nothing you run here is recorded.

Scratch console

A scratch console with the scientific stack (pandas, numpy, scikit-learn). Runs on the server — no network, resource-limited and measured.

Output appears here.

My notes on this module

Loading your notes...

Where next?

Continue

Keeping Softmax in Its Range →

Attention · P10

This module unlocks

Multi-Head Attention — TransformerLab