Batch Normalization
BatchNorm is a compact paper for learning how activation statistics, optimization stability, and architecture design interact. It is a good bridge between practical deep learning and the math of normalization.
Reading focus: Why unstable activation distributions can make optimization harder. How learned scale and shift keep normalization from removing capacity. Why training-time statistics and inference-time statistics differ.
ICML 2015. Ioffe and Szegedy. 35 min read, medium difficulty.