Guided pathThis is part of Understand Transformers & BERTBack to the path

Keeping Softmax in Its Range

P10.attention.04 · Audience: guest, it-ml, language-pro · Prerequisites: Multi-Head Attention

Customise the corpus these labs use

A softmax turns scores into a distribution, and what kind of distribution it produces depends almost entirely on how large the scores are. Small scores give a spread of attention a model can learn from; large ones give a single winner and a gradient close to zero. This module reproduces the failure on this platform's own corpora — the same code, seed and sentence, breaking as the corpus grows — then takes it apart into the three guards that keep a softmax in its range, one switch each.

Step 1 / 5 — The Failure, Reproduced
🗣️ From a linguist's perspective: a scale that changed under the reader
A dictionary definition read aloud in a whisper and in a shout says the same thing, but a listener who judges importance by volume will hear only the shout. Attention judges importance by the size of a score, so the scale of what it reads matters as much as the content.
ⓘ Concept: The same recipe on three corpora
The panel runs one attention head on one sentence with the projections drawn at 1/√d_model and the scores divided by √d_k — two of the three guards — but reads the position-aware sum exactly as the corpus produces it. Pick the eight-sentence demo corpus, then the two curated ones. Nothing in the code changes; only the corpus the embeddings were learned from. Watch the input's spread, the scores' spread and the mean row-max climb together until every row puts almost all its weight on one token.
weights=softmax ⁣((XWQ)(XWK)⊤dk),W∼N ⁣(0,1dmodel)\text{weights} = \mathrm{softmax}\!\left(\frac{(X W_Q)(X W_K)^{\top}}{\sqrt{d_k}}\right),\qquad W \sim \mathcal{N}\!\left(0, \tfrac{1}{d_{\text{model}}}\right)

Why it matters — This is the recipe every attention surface here used until the input was normalised: correct for small inputs, and silently wrong as soon as the corpus grew.

The same recipe, three corpora

Reads /api/attention/scale-guards — loading

📚 Go further

Check your understanding

Self-check

If the entries of the attention input double in size, what happens to the attention scores?

Self-check

Why does dividing the scores by √d_k not rescue attention on the raw sum of a large corpus?

Self-check

Why is BERT unaffected by the scale of its embeddings?

Ask the mentor about this module

Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.

Images, PDF or text. Kept on this device only.
Keeping your files on this device

Off by default. The mentor always gets your file; this only decides whether your own copy stays here. Copies live in this browser only - they do not follow you to another device, and clearing site data removes them.

Ctrl/Cmd + Enter to send

Try it yourself

A scratch console for this page's ideas — ungraded, nothing you run here is recorded.

Scratch console

A scratch console with the scientific stack (pandas, numpy, scikit-learn). Runs on the server — no network, resource-limited and measured.

Output appears here.

My notes on this module

Loading your notes...

Where next?

Continue

Encoder Block →

Transformer Architecture · P10

Keeping Softmax in Its Range — TransformerLab