Keeping Softmax in Its Range
P10.attention.04 · Audience: guest, it-ml, language-pro · Prerequisites: Multi-Head Attention
A softmax turns scores into a distribution, and what kind of distribution it produces depends almost entirely on how large the scores are. Small scores give a spread of attention a model can learn from; large ones give a single winner and a gradient close to zero. This module reproduces the failure on this platform's own corpora — the same code, seed and sentence, breaking as the corpus grows — then takes it apart into the three guards that keep a softmax in its range, one switch each.
🗣️ From a linguist's perspective: a scale that changed under the reader
ⓘ Concept: The same recipe on three corpora
Why it matters — This is the recipe every attention surface here used until the input was normalised: correct for small inputs, and silently wrong as soon as the corpus grew.
The same recipe, three corpora
Reads /api/attention/scale-guards — loading
📚 Go further
- Vaswani et al. 2017, §3.2.1 — Attention Is All You Need
- Ba, Kiros, Hinton 2016 — Layer Normalization
- Devlin et al. 2019, §3 — BERT
Check your understanding
Self-check
If the entries of the attention input double in size, what happens to the attention scores?
Self-check
Why does dividing the scores by √d_k not rescue attention on the raw sum of a large corpus?
Self-check
Why is BERT unaffected by the scale of its embeddings?
Ask the mentor about this module
Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.
Keeping your files on this device
Off by default. The mentor always gets your file; this only decides whether your own copy stays here. Copies live in this browser only - they do not follow you to another device, and clearing site data removes them.
Try it yourself
A scratch console for this page's ideas — ungraded, nothing you run here is recorded.
Scratch console
A scratch console with the scientific stack (pandas, numpy, scikit-learn). Runs on the server — no network, resource-limited and measured.
Output appears here.
My notes on this module
Loading your notes...
Where next?
Go up a level