Labs · Inside the block

Tokens expanded and projected

The feed-forward sublayer one position at a time: widen to d_ff, activate, narrow back — and what the activation and the expansion each buy.

FFN(x)=W2 ϕ(W1x+b1)+b2\mathrm{FFN}(x) = W_2\, \phi(W_1 x + b_1) + b_24 rows · d_model = 4 → d_ff = 16 → 4

BERT's activation: smooth, and slightly negative just below zero.

Widen, activate, narrow — one row at a time

Reads /api/labs/ffn — loading

Computing…

Tokens expanded and projected — Labs — TransformerLab