Adam: A Method for Stochastic Optimization

Adam is the optimizer paper learners meet constantly in code before they understand it. The first pass is approachable: it combines momentum-like gradient averages with squared-gradient adaptation.

Reading focus: How first and second moment estimates change parameter updates. Why bias correction matters early in training. What tradeoffs make Adam convenient but not automatically best.

ICLR 2015. Kingma and Ba. 30 min read, very easy difficulty.