Labs · Inside the block

Rows normalised

Each row to mean 0 and variance 1 on its own, then γ and β; where ε matters, and what batch norm would have done to the same matrix instead.

LN(x)=γ⊙x−μrowσrow2+ε+β\mathrm{LN}(x) = \gamma \odot \frac{x - \mu_{\text{row}}}{\sqrt{\sigma^2_{\text{row}} + \varepsilon}} + \beta4 rows × 8 features

Four rows a hundredfold apart in scale: each comes out at mean 0, variance 1.

Before and after, row by row

Reads /api/labs/layernorm-rows — loading

Computing…

Rows normalised — Labs — TransformerLab