Paper Lineage
Esc
evaluationIntroduced by Chinchilla · 2022

Compute-optimal training

For a fixed budget, grow parameters and training tokens together; most big LMs were undertrained.

Drafted by AI · not yet reviewed

How this idea evolved

Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.

  1. 2020

    Loss falls as a power law in model size, data and compute.

    Cites · not yet reviewedcited 12× · §Introduction
    “Following Kaplan et al. 2020 and the training setup of GPT-3 (Brown et al. 2020), many of the recently trained large models have been trained for approximately 300 billion tokens (Table 1), in line with the approach of predominantly increasing model size when increasing compute.”
    From Chinchilla · §Introduction
  2. 2022

    For a fixed budget, grow parameters and training tokens together; most big LMs were undertrained.

Papers using this