Training Compute-Optimal Large Language Models
Chinchilla is the scaling paper that made data budget feel as important as parameter count. It is a good daily pick for learning why a smaller model trained on more tokens can beat a larger undertrained one.
Reading focus: Why compute-optimal training balances model size and token count. How scaling laws become practical training-budget decisions. Why undertrained large models can waste compute despite looking impressive.
arXiv 2022. Hoffmann et al.. 45 min read, medium difficulty.