Paper Lineage
Esc
AnalysisMar 2022arXiv 2203.15556cs.CL

Training Compute-Optimal Large Language Models

Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch and 19 others

We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant.

From the abstract

Built on

9 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (9)

  • Gopher2021 · cited 16×
    “These include both dense transformer models (Brown et al. 2020; Lieber et al. 2021; Smith et al. 2022; Rae et al. 2021; Thoppilan et al. 2022) and mixture-of-expert (MoE) models (Du et al. 2021; Fedus et al. 2021; Zoph e…”
    From this paper · §Related Work
  • Kaplan scaling laws2020 · cited 12×
    “Following Kaplan et al. 2020 and the training setup of GPT-3 (Brown et al. 2020), many of the recently trained large models have been trained for approximately 300 billion tokens (Table 1), in line with the approach of p…”
    From this paper · §Introduction
  • LaMDA2022 · cited 4×
    “These include both dense transformer models (Brown et al. 2020; Lieber et al. 2021; Smith et al. 2022; Rae et al. 2021; Thoppilan et al. 2022) and mixture-of-expert (MoE) models (Du et al. 2021; Fedus et al. 2021; Zoph e…”
    From this paper · §Related Work
  • GPT-32020 · cited 3×
    “Following Kaplan et al. 2020 and the training setup of GPT-3 (Brown et al. 2020), many of the recently trained large models have been trained for approximately 300 billion tokens (Table 1), in line with the approach of p…”
    From this paper · §Introduction
Show 5 more
  • Megatron-Turing NLG2022 · cited 3×
    “These include both dense transformer models (Brown et al. 2020; Lieber et al. 2021; Smith et al. 2022; Rae et al. 2021; Thoppilan et al. 2022) and mixture-of-expert (MoE) models (Du et al. 2021; Fedus et al. 2021; Zoph e…”
    From this paper · §Related Work
  • MMLU2020 · cited 2×
    “We thus place more emphasis on other tasks for which leakage is less of a concern, such as MMLU (Hendrycks et al. 2020) and BIG-bench (BIG-bench collaboration 2021) along with various closed-book question answering and c…”
    From this paper · §Chinchilla
  • Switch Transformer2021 · cited 2×
    “These include both dense transformer models (Brown et al. 2020; Lieber et al. 2021; Smith et al. 2022; Rae et al. 2021; Thoppilan et al. 2022) and mixture-of-expert (MoE) models (Du et al. 2021; Fedus et al. 2021; Zoph e…”
    From this paper · §Related Work
  • GLaM2021 · cited 2×
    “These include both dense transformer models (Brown et al. 2020; Lieber et al. 2021; Smith et al. 2022; Rae et al. 2021; Thoppilan et al. 2022) and mixture-of-expert (MoE) models (Du et al. 2021; Fedus et al. 2021; Zoph e…”
    From this paper · §Related Work
  • BIG-bench2022 · cited 2×
    “We thus place more emphasis on other tasks for which leakage is less of a concern, such as MMLU (Hendrycks et al. 2020) and BIG-bench (BIG-bench collaboration 2021) along with various closed-book question answering and c…”
    From this paper · §Chinchilla

Led to

  • PaLM2022 · cited 5×
    “The most powerful of these post-GPT-3 models are GLaM (Du et al. 2021), Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), Megatron–Turing NLG (Smith et al. 2022), and LaMDA (Thoppilan et al. 2022), all of whic…”
    From PaLM · §Introduction
  • Flamingo2022 · cited 5×, 1 in Method
    “We perform experiments across three models sizes, building on the 1.4B, 7B, and 70B parameter Chinchilla models [42]; calling them respectively Flamingo-3B, Flamingo-9B and Flamingo-80B.”
    From Flamingo · §Approach
  • OPT2022 · cited 4×
    “These scaling gains come not only from growing the total number of parameters in the models, but also the amount and quality of pre-training data Liu et al. 2019b; Hoffmann et al. 2022.”
    From OPT · §Related Work
  • BIG-bench2022 · cited 2×
    “BIG-bench has also been partially evaluated on other models including Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), and T0 (Sanh et al. 2022).”
    From BIG-bench · §What is in BIG-bench?
  • Emergent abilities2022 · cited 6×
    “In many cases, the effect of scale on performance can often be methodologically predicted via scaling laws—for example, scaling curves for cross-entropy loss have been shown to empirically span more than seven orders of…”
    From Emergent abilities · §Introduction
  • U-PaLM2022 · cited 8×
    “To this end, large language models not only continue to improve as we scale in terms of data or computational budget (Hoffmann et al. 2022; Kaplan et al. 2020) but also acquire new abilities (Wei et al. 2022a).”
    From U-PaLM · §Related Work
Abstract

We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant. By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled. We test this hypothesis by training a predicted compute-optimal model, Chinchilla, that uses the same compute budget as Gopher but with 70B parameters and 4$\times$ more more data. Chinchilla uniformly and significantly outperforms Gopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on a large range of downstream evaluation tasks. This also means that Chinchilla uses substantially less compute for fine-tuning and inference, greatly facilitating downstream usage. As a highlight, Chinchilla reaches a state-of-the-art average accuracy of 67.5% on the MMLU benchmark, greater than a 7% improvement over Gopher.