Worked Example — one BERT by hand

A track of P10 · Generative AI.

One fixed 8-dimensional configuration carried from raw corpus to a weight update, every matrix named and every number the platform's own — the reference trace to check an intuition against.

Every other track in this pillar teaches one stage and lets you turn its knobs. That is the right shape for learning what attention does, and the wrong shape for answering a different question: what actually happens, in order, to one sentence, with numbers you can check?

This track answers that. It fixes a single configuration — the smallest the platform will accept — and carries it from three sentences of raw text all the way to a weight that has moved. Nothing is re-parameterised between steps. The matrix leaving one section is visibly the matrix entering the next, and every number on the page is generated by the same code the rest of the pillar runs, not transcribed into prose that quietly goes stale.

It opens with a table of notation, and that table is the point as much as the trace is. Attention has a naming problem: the same object is a context vector here and an attention output there, the feed-forward matrices are W_1 and W_2 in one paper and ff_in and ff_out in the next. Names that move between explanations make a derivation impossible to follow. So the aliases are stated once, at the start, and one name is used from then on.

Read it in order the first time — the four modules are one continuous argument. Afterwards it is a reference: come back to check a shape, settle what W_O does, or find out whether the thing you are reading about is the same object under a different name.

The third module closes two gaps the pillar leaves open elsewhere: what separates an encoder block from a decoder block, and what BERT trains besides masked words — including why the [CLS] token exists at all, and why it is still a poor sentence summary despite it. The fourth counts what the trace trained: it defines a parameter, adds this configuration's 1,833 up by hand, and scales the same arithmetic until it reproduces the platform's own trained BERT — and the headline numbers on model cards.

Corpus → X_postokens, E lookup, positional encodingAttentionprojections, head slices, softmax, contextBlock outputresidual, LayerNorm, feed-forward, againLoss → updatelogits, cross-entropy, gradients, one stepThe countthe parameter inventory, the formula, both engines reconciled
One configuration, five stages, no re-parameterising between them — the matrix leaving each stage is the matrix entering the next.
Worked Example — one BERT by hand — TransformerLab