Sentence Embeddings
P10.transformer-architecture.02 · Audience: guest, it-ml, language-pro · Prerequisites: Encoder Block
The encoder block leaves us with one contextualised vector per token, but sentences have different lengths, and comparing, searching, or clustering them requires a single fixed-size vector per sentence. Pooling is the step that collapses the token rows into that one vector, and there is more than one defensible way to do it. This capstone module walks through the three classic strategies — take the first token, average them all, or keep each dimension's maximum — and then puts the result to work: every corpus sentence runs through the full pipeline, and the lab ranks them all by similarity to the query sentence you select.
🗣️ From a linguist's perspective: from words to the sentence
ⓘ Concept: From token representations to sentence vectors
Why it matters — Search, clustering, and semantic comparison all need ONE fixed-size vector per text — pooling is the bridge from token space to sentence space.
Reference: Reimers & Gurevych 2019 — Sentence-BERT
Computing…
Open in Labs: Pooling with and without the mask
🎓 Interview deep-dive
Adapting a pre-trained model. BERT-style: fine-tune with a task head on [CLS] (or mean-pool for sentence tasks — SBERT). GPT-style: prompt with in-context/few-shot examples, no weight updates. Between them: parameter-efficient fine-tuning (LoRA/PEFT — train tiny adapters), instruction tuning, and RAG (retrieve with pooled embeddings exactly like step 6, then generate grounded on the hits).
📚 Go further
- Reimers & Gurevych 2019 — Sentence-BERT
- Jay Alammar — The Illustrated BERT
- Sentence-Transformers — pretrained models
- Devlin et al. 2019 — BERT
- Hugging Face — NLP course, ch. 5–6
Check your understanding
Self-check
What is the encoder-output shape for 3 tokens with d_model = 8, and after pooling?
Self-check
Which token does BERT use as the sentence embedding?
Self-check
Which pooling strategy is most sensitive to a single highly-activated token?
Self-check
Two sentence vectors have cosine similarity 0.95 — similar or dissimilar?
Self-check
What advantage does mean pooling have over CLS without fine-tuning?
Ask the mentor about this module
Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.
Keeping your files on this device
Off by default. The mentor always gets your file; this only decides whether your own copy stays here. Copies live in this browser only - they do not follow you to another device, and clearing site data removes them.
Try it yourself
A scratch console for this page's ideas — ungraded, nothing you run here is recorded.
Scratch console
A scratch console with the scientific stack (pandas, numpy, scikit-learn). Runs on the server — no network, resource-limited and measured.
Output appears here.
My notes on this module
Loading your notes...
Where next?
This module unlocks
Go up a level