What is Attention?
P10.attention.01 · Audience: guest, it-ml, language-pro
By this point in the pipeline every word of the sentence has become a vector — a list of numbers standing in for its meaning — but that vector is context-blind: the row for bank is exactly the same in "river bank" and in "savings bank". A sentence only makes sense because its words clarify one another, so the model needs a way to let each word's representation absorb meaning from the words that disambiguate it. That operation is Attention, and before computing a single number we look at what it is for, through the eyes of a linguist, a mathematician, a statistician and a programmer in turn.
By the end of the text-representation and positional-encoding modules, every word of the sentence has become a vector that encodes what the word means (its SVD embedding) and where it stands (its positional encoding). One more step comes before attention: a LayerNorm puts every row on the same scale. Follow one word through those three boxes below. The whole sentence is one matrix — one row per word, d_model numbers per row.
Computing…
But every row is still context-blind: the vector for bank is identical in "river bank" and "savings bank". The question Attention answers is:
For each word, which other words matter right now — and how do we mix their meanings into its representation?
Open in Labs: The attention matrix
One honest note about this page: nothing on this page trains. The attention maps it draws come from seeded stand-in weights, so you can study the mechanism in isolation. The weights a trained model would use are what training produces — and the worked example shows that happening: batch by batch, not one sentence at a time. And one module on, the scaled dot-product demo can now draw the trained maps themselves beside its stand-ins, straight from the model store. For the same computation with several heads over a three-layer stack, open the attention demo.
Ask the mentor about this module
Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.
Keeping your files on this device
Off by default. The mentor always gets your file; this only decides whether your own copy stays here. Copies live in this browser only - they do not follow you to another device, and clearing site data removes them.
Try it yourself
A scratch console for this page's ideas — ungraded, nothing you run here is recorded.
Scratch console
A scratch console with the scientific stack (pandas, numpy, scikit-learn). Runs on the server — no network, resource-limited and measured.
Output appears here.
My notes on this module
Loading your notes...
Where next?
Go up a level