Labs · Training

A gradient down twelve layers

What the backward pass still carries at each depth — with the residual path, without it, and with LayerNorm before or after.

hl+1=hl+ϕ(hlWl+bl)h_{l+1} = h_l + \phi(h_l W_l + b_l)12 layers × width 8 · a batch of 4

Twelve tanh layers and nothing else: the first layer receives under 1 % of the signal.

What the backward pass still carries at each depth

Reads /api/labs/gradient-depth — loading

Computing…

A gradient down twelve layers — Labs — TransformerLab