Labs · Training

SGD, momentum and Adam

The same gradient taken three ways across a drawn surface: plain descent, descent with velocity, and a per-coordinate step — and what the learning rate does to each.

θ←θ−η m^v^+ϵ\theta \leftarrow \theta - \eta\, \frac{\hat m}{\sqrt{\hat v} + \epsilon}a 2-D surface · 60 steps

Curvature 20× steeper along y: SGD zig-zags, momentum overshoots, Adam adapts per axis.

Three routes down the same surface

Reads /api/labs/optimizers — loading

Computing…

SGD, momentum and Adam — Labs — TransformerLab