WordPiece Tokenization
P10.text-representation.05 · Audience: guest, it-ml, language-pro · Prerequisites: BPE Tokenization
Merging the most frequent pair, as BPE does, is not the only reasonable way to grow a vocabulary: a pair of pieces can be rare yet so exclusive that they almost never appear apart, and that exclusivity is strong evidence of a real linguistic unit. WordPiece, the tokenizer behind BERT, keeps BPE's merge loop but ranks candidate pairs by exactly that specificity instead of raw frequency. This module traces WordPiece stage by stage on the same corpus as the previous module, so you can see precisely where the two tokenizers diverge and why BERT and GPT split the same word differently.
🗣️ From a linguist's perspective: Specificity over raw frequency
ⓘ Concept: Three differences from BPE
WordPiece runs the same iterative merge loop as BPE, but:
| BPE | WordPiece | |
|---|---|---|
| Word boundary | </w> suffix | ## prefix on non-initial tokens |
| Merge score | freq(AB) | freq(AB) / (freq(A) × freq(B)) |
| Inference | replay merge rules in order | longest-match-first in the vocabulary |
The ## prefix inverts the mental model: BPE marks where a word ends; WordPiece marks this token is glued to whatever came before it.
Why it matters — Because the score divides by freq(A) × freq(B), a pair whose parts appear only together scores 1.0 — the maximum — even if it occurs once. BPE would never pick it. That single change is why BERT and GPT produce different tokens for the same word.
Reference: Schuster & Nakamura 2012 — Japanese and Korean Voice Search (WordPiece)
Used by: BERT, DistilBERT, mBERT. The capstone compares WordPiece with BPE live.
Reading the corpus…
Open in Labs: BPE against WordPiece
Check your understanding
Self-check
WordPiece and BPE run the same merge loop. Name the one quantity they rank candidate pairs by that differs.
Self-check
A pair appears only once, but its two tokens only ever appear together. What is its WordPiece score, and will BPE merge it early?
Self-check
BPE stores its ordered merge rules; WordPiece does not. What does WordPiece use at inference time instead?
Ask the mentor about this module
Ask a question about this content. The mentor explains and grounds its answer in what you are studying; asking is recorded as a learning signal, not a grade.
Keeping your files on this device
Off by default. The mentor always gets your file; this only decides whether your own copy stays here. Copies live in this browser only - they do not follow you to another device, and clearing site data removes them.
Try it yourself
A scratch console for this page's ideas — ungraded, nothing you run here is recorded.
Scratch console
A scratch console with the scientific stack (pandas, numpy, scikit-learn). Runs on the server — no network, resource-limited and measured.
Output appears here.
My notes on this module
Loading your notes...
Where next?
Later in Text Representation
Go up a level