Guided pathThis is part of Understand Transformers & BERTBack to the path

What is Attention?

P10.attention.01 · Audience: guest, it-ml, language-pro

Customise the corpus these labs use

By this point in the pipeline every word of the sentence has become a vector — a list of numbers standing in for its meaning — but that vector is context-blind: the row for bank is exactly the same in "river bank" and in "savings bank". A sentence only makes sense because its words clarify one another, so the model needs a way to let each word's representation absorb meaning from the words that disambiguate it. That operation is Attention, and before computing a single number we look at what it is for, through the eyes of a linguist, a mathematician, a statistician and a programmer in turn.

Step 1 / 3 — Where we are in the pipeline

By the end of the text-representation and positional-encoding modules, every word of the sentence has become a vector that encodes what the word means (its SVD embedding) and where it stands (its positional encoding). One more step comes before attention: a LayerNorm puts every row on the same scale. Follow one word through those three boxes below. The whole sentence is one matrix — one row per word, d_model numbers per row.

Computing…

But every row is still context-blind: the vector for bank is identical in "river bank" and "savings bank". The question Attention answers is:

For each word, which other words matter right now — and how do we mix their meanings into its representation?

Open in Labs: The attention matrix

One honest note about this page: nothing on this page trains. The attention maps it draws come from seeded stand-in weights, so you can study the mechanism in isolation. The weights a trained model would use are what training produces — and the worked example shows that happening: batch by batch, not one sentence at a time. And one module on, the scaled dot-product demo can now draw the trained maps themselves beside its stand-ins, straight from the model store. For the same computation with several heads over a three-layer stack, open the attention demo.

Ask the mentor about this module

Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.

Images, PDF or text. Kept on this device only.
Keeping your files on this device

Off by default. The mentor always gets your file; this only decides whether your own copy stays here. Copies live in this browser only - they do not follow you to another device, and clearing site data removes them.

Ctrl/Cmd + Enter to send

Try it yourself

A scratch console for this page's ideas — ungraded, nothing you run here is recorded.

Scratch console

A scratch console with the scientific stack (pandas, numpy, scikit-learn). Runs on the server — no network, resource-limited and measured.

Output appears here.

My notes on this module

Loading your notes...

What is Attention? — TransformerLab