Labs · Training
A gradient down twelve layers
What the backward pass still carries at each depth — with the residual path, without it, and with LayerNorm before or after.
12 layers × width 8 · a batch of 4
Twelve tanh layers and nothing else: the first layer receives under 1 % of the signal.
What the backward pass still carries at each depth
Reads /api/labs/gradient-depth — loading
Computing…