Labs
Take one mechanism and play with it. Every lab runs on small test data you can edit — sized so that every number fits on the screen and changing an input visibly changes the result — and every number is computed by the same functions the modules use, so a lab and a module never disagree.
Concept labs
One mechanism, on test data you can edit, sized to show everything it does.
Scores and attention
How a row of numbers becomes a row of weights.
The attention matrix
Each query scores every key, the scores become weights, and the weights mix the values.
6 tokens · d_k = d_v = 4
OpenSoftmax
Logits become probabilities; temperature, masking and scale decide how peaked they are and whether a gradient survives.
8 logits
OpenHeads split and mixed
One attention split into heads by slicing the projections, each head's map on its own, the contexts concatenated and mixed back by W_O.
6 tokens · d_model = 8 · 2 heads of d_k = 4
Open
Vectors
What a dot product, a matrix product and a pooled vector actually compute.
Dot product and cosinefrom a module
Two vectors, one number: how much they point the same way, with and without their lengths.
OpenMatrix multiplicationfrom a module
Every cell of the product is one dot product — rows of A against columns of B.
OpenPooling with and without the mask
One sentence vector from a batch's token rows — [CLS], mean or max — and what the padding rows do to it unless the mask keeps them out.
8 rows × 4 dims · 2 of them padding
OpenDot product against cosine
The same vectors under both measures, whose neighbour changes, and anisotropy — a shared offset that makes every cosine high until the mean is subtracted.
8 vectors × 4 dims · two clusters
Open
Inside the block
The rest of one encoder block, one sublayer at a time.
Dropoutfrom a module
Zero a random share of activations while training, and scale the rest so the sum holds.
OpenLayer normalisationfrom a module
Centre and scale one row to zero mean and unit variance, then let γ and β undo what training wants undone.
OpenThe sinusoidal tablefrom a module
Position as a fingerprint of sines and cosines at geometrically spaced frequencies.
OpenRoPE and ALiBifrom a module
Position as a rotation of the query–key pair, or as a penalty on distance — the two modern answers.
OpenPositions encoded and compared
The sinusoidal table and why PE·PEᵀ depends only on the offset; the same property as a rotation (RoPE) and as a penalty on distance (ALiBi).
32 positions × d_model = 16
OpenRows normalised
Each row to mean 0 and variance 1 on its own, then γ and β; where ε matters, and what batch norm would have done to the same matrix instead.
4 rows × 8 features
OpenTokens expanded and projected
The feed-forward sublayer one position at a time: widen to d_ff, activate, narrow back — and what the activation and the expansion each buy.
4 rows · d_model = 4 → d_ff = 16 → 4
Open
Training
What happens across one step, many steps, and many layers.
The perceptronfrom a module
A weighted sum, a threshold, and the smallest unit that can learn a line.
OpenCross-entropy against MSE
The same batch of predictions scored two ways — how hard each loss punishes a confident mistake, what a batch mean hides, and which positions a mask lets count.
5 classes · a batch of 4
OpenForward and backward
A small MLP run forward, its loss sent back through the chain rule, every gradient checked against the definition of a derivative, and one step taken.
3 → 4 → 4 → 3 · a batch of 4
OpenSGD, momentum and Adam
The same gradient taken three ways across a drawn surface: plain descent, descent with velocity, and a per-coordinate step — and what the learning rate does to each.
a 2-D surface · 60 steps
OpenA gradient down twelve layers
What the backward pass still carries at each depth — with the residual path, without it, and with LayerNorm before or after.
12 layers × width 8 · a batch of 4
Open
Size and tokens
How big the model is, and what the vocabulary under it is made of.
Parameters, FLOPs and memory
A model's size from its dials — every parameter counted by the platform's own derivation, the FLOPs a token costs, and what training it holds in memory.
this platform's model, up to a 7B decoder
OpenOne distribution, one token
Greedy, temperature, top-k and top-p on one next-token distribution, with a thousand seeded draws to show what each keeps and what each throws away.
12 candidates · 1,000 draws
OpenBPE against WordPiece
Two merge rules on one editable word table — the most frequent pair against the highest-scoring one — and how each then splits a word it has never seen.
10 words · 20 merges
Open
Worked examples
One small model computed end to end, every number on the page.
The traced BERT — 3 sentences, then 64
One small BERT computed end to end: notation, forward pass, loss and backward, the parameter count.
OpenWord2vec, step by stepplanned
One skip-gram training pair traced through lookup, softmax, error and both gradient paths.
Planned — #186
Workflows
The concepts connected, on a real corpus.
Real cases
The platform applied to your own text or to a production lifecycle.
Your own text
The whole pipeline on a text you paste — with the lens that produced each result named.
OpenProduction case — the income model
A tabular model, packaged to monitored.
OpenProduction case — TinyBERT
A transformer, packaged to monitored.
OpenThe corpus-scale analysisplanned
Where signal emerges from noise as the corpus grows — the ladder as a story.
Planned — #1925