Learned Embeddings
P10.text-representation.06 · Audience: guest, it-ml, language-pro · Prerequisites: Embeddings
The embedding table we built from co-occurrence counts is frozen: every word keeps one fixed vector forever, whatever sentence it appears in, and no feedback ever improves it. Modern models take the opposite approach and learn the table by training on a task, hiding a word and then nudging every weight that failed to predict it. This module traces one such BERT-style training step end to end with real numbers, from masking a token to the gradient that rewrites the masked word's row of the embedding table.
🗣️ From a linguist's perspective: why context changes meaning
ⓘ Concept: Two paradigms
Why it matters — The leap from static (P10.text-representation.03) to learned + contextual is the leap from LSA-era NLP to BERT and GPT.
Reference: Devlin et al. 2018 — BERT
Computing…
🎓 Interview deep-dive
BERT vs GPT pretraining. BERT masks 15% of tokens (with the 80/10/10 trick: 80% [MASK], 10% random token, 10% unchanged) and predicts them bidirectionally; GPT predicts every next token left-to-right. GPT's objective yields a loss signal at every position (dense), BERT's only at masked ones (sparse) — one reason autoregressive scaling proved so effective. This page's variant feeds the true token (CBOW-with-attention) so the gradient reaches the true row directly — same dynamic, no special-token bookkeeping.
📚 Go further
- Devlin et al. 2018 — BERT
- Jay Alammar — The Illustrated BERT
- Andrej Karpathy — Let's build GPT
Check your understanding
Self-check
What three components sum to form the BERT input embedding?
Self-check
Why does dL/dlogits sum to exactly zero?
Self-check
What happens to the update if you raise the learning rate from 0.01 to 0.5?
Ask the mentor about this module
Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.
Keeping your files on this device
Off by default. The mentor always gets your file; this only decides whether your own copy stays here. Copies live in this browser only - they do not follow you to another device, and clearing site data removes them.
Try it yourself
A scratch console for this page's ideas — ungraded, nothing you run here is recorded.
Scratch console
A scratch console with the scientific stack (pandas, numpy, scikit-learn). Runs on the server — no network, resource-limited and measured.
Output appears here.
My notes on this module
Loading your notes...
Where next?
Go up a level