Scaling Laws for Neural Language Models

The scaling-laws paper turns model size, data, and compute into a quantitative tradeoff. It is useful when you want to reason about progress curves instead of treating bigger training runs as folklore.

Reading focus: How loss trends with parameter count, dataset size, and compute budget. Why log-log plots are the natural language of empirical scaling. Where simple laws help and where newer regimes can break them.

arXiv 2020. Kaplan et al.. 50 min read, medium difficulty.