evaluationIntroduced by Chinchilla · 2022
Compute-optimal training
For a fixed budget, grow parameters and training tokens together; most big LMs were undertrained.
Drafted by AI · not yet reviewed
How this idea evolved
Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.
- 2020
Loss falls as a power law in model size, data and compute.
Cites · not yet reviewedcited 12× · §Introduction“Following Kaplan et al. 2020 and the training setup of GPT-3 (Brown et al. 2020), many of the recently trained large models have been trained for approximately 300 billion tokens (Table 1), in line with the approach of predominantly increasing model size when increasing compute.”
From Chinchilla · §Introduction - 2022
For a fixed budget, grow parameters and training tokens together; most big LMs were undertrained.