Scaling Laws for Neural Language Models
The scaling-laws paper turns model size, data, and compute into a quantitative tradeoff. It is useful when you want to reason about progress curves instead of treating bigger training runs as folklore.
Reading focus: How loss trends with parameter count, dataset size, and compute budget. Why log-log plots are the natural language of empirical scaling. Where simple laws help and where newer regimes can break them.
arXiv 2020. Kaplan et al.. 50 min read, medium difficulty.