Embeddings
P10.text-representation.03 · Audience: guest, it-ml, language-pro · Prerequisites: Tokenization
A token ID is only a shelf mark: the number 17 says nothing about what the word means, and a model can compute nothing useful with it. What we need instead is a representation in which similar words receive similar numbers, so that meaning becomes something we can measure. This module builds exactly that from the corpus itself: it counts which words keep company with which, compresses those counts into dense vectors with truncated SVD, and looks up your sentence's rows in the resulting embedding table. Every number you will see here is derived from corpus statistics, never from random initialisation.
🗣️ From a linguist's perspective: the meaning dictionary
ⓘ Concept: The table
Why it matters — This table is the model's entire lexical knowledge — GloVe shipped exactly such a matrix for 400k words; every LLM still starts with one.
Reference: Pennington et al. 2014 — GloVe
Computing…
📚 Go further
- Jay Alammar — The Illustrated Word2Vec
- Pennington et al. — GloVe project page
- Levy & Goldberg 2014 — Implicit matrix factorization
Check your understanding
Self-check
Where do the numbers in the embedding table come from on this page?
Self-check
What does a larger window size change in the co-occurrence matrix?
Self-check
Why is the embedding lookup so cheap?
Ask the mentor about this module
Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.
Keeping your files on this device
Off by default. The mentor always gets your file; this only decides whether your own copy stays here. Copies live in this browser only - they do not follow you to another device, and clearing site data removes them.
Try it yourself
A scratch console for this page's ideas — ungraded, nothing you run here is recorded.
Scratch console
A scratch console with the scientific stack (pandas, numpy, scikit-learn). Runs on the server — no network, resource-limited and measured.
Output appears here.
My notes on this module
Loading your notes...
Where next?
Later in Text Representation
This module unlocks
Go up a level