Training Compute-Optimal Large Language Models
We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant.
Also cited · not yet reviewed (9)
- Gopher2021 · cited 16דThese include both dense transformer models (Brown et al. 2020; Lieber et al. 2021; Smith et al. 2022; Rae et al. 2021; Thoppilan et al. 2022) and mixture-of-expert (MoE) models (Du et al. 2021; Fedus et al. 2021; Zoph e…”From this paper · §Related Work
- Kaplan scaling laws2020 · cited 12דFollowing Kaplan et al. 2020 and the training setup of GPT-3 (Brown et al. 2020), many of the recently trained large models have been trained for approximately 300 billion tokens (Table 1), in line with the approach of p…”From this paper · §Introduction
- LaMDA2022 · cited 4דThese include both dense transformer models (Brown et al. 2020; Lieber et al. 2021; Smith et al. 2022; Rae et al. 2021; Thoppilan et al. 2022) and mixture-of-expert (MoE) models (Du et al. 2021; Fedus et al. 2021; Zoph e…”From this paper · §Related Work
- GPT-32020 · cited 3דFollowing Kaplan et al. 2020 and the training setup of GPT-3 (Brown et al. 2020), many of the recently trained large models have been trained for approximately 300 billion tokens (Table 1), in line with the approach of p…”From this paper · §Introduction
Show 5 more
- Megatron-Turing NLG2022 · cited 3דThese include both dense transformer models (Brown et al. 2020; Lieber et al. 2021; Smith et al. 2022; Rae et al. 2021; Thoppilan et al. 2022) and mixture-of-expert (MoE) models (Du et al. 2021; Fedus et al. 2021; Zoph e…”From this paper · §Related Work
- MMLU2020 · cited 2דWe thus place more emphasis on other tasks for which leakage is less of a concern, such as MMLU (Hendrycks et al. 2020) and BIG-bench (BIG-bench collaboration 2021) along with various closed-book question answering and c…”From this paper · §Chinchilla
- Switch Transformer2021 · cited 2דThese include both dense transformer models (Brown et al. 2020; Lieber et al. 2021; Smith et al. 2022; Rae et al. 2021; Thoppilan et al. 2022) and mixture-of-expert (MoE) models (Du et al. 2021; Fedus et al. 2021; Zoph e…”From this paper · §Related Work
- GLaM2021 · cited 2דThese include both dense transformer models (Brown et al. 2020; Lieber et al. 2021; Smith et al. 2022; Rae et al. 2021; Thoppilan et al. 2022) and mixture-of-expert (MoE) models (Du et al. 2021; Fedus et al. 2021; Zoph e…”From this paper · §Related Work
- BIG-bench2022 · cited 2דWe thus place more emphasis on other tasks for which leakage is less of a concern, such as MMLU (Hendrycks et al. 2020) and BIG-bench (BIG-bench collaboration 2021) along with various closed-book question answering and c…”From this paper · §Chinchilla
Led to
- PaLM2022 · cited 5דThe most powerful of these post-GPT-3 models are GLaM (Du et al. 2021), Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), Megatron–Turing NLG (Smith et al. 2022), and LaMDA (Thoppilan et al. 2022), all of whic…”From PaLM · §Introduction
- Flamingo2022 · cited 5×, 1 in Method“We perform experiments across three models sizes, building on the 1.4B, 7B, and 70B parameter Chinchilla models [42]; calling them respectively Flamingo-3B, Flamingo-9B and Flamingo-80B.”From Flamingo · §Approach
- OPT2022 · cited 4דThese scaling gains come not only from growing the total number of parameters in the models, but also the amount and quality of pre-training data Liu et al. 2019b; Hoffmann et al. 2022.”From OPT · §Related Work
- BIG-bench2022 · cited 2דBIG-bench has also been partially evaluated on other models including Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), and T0 (Sanh et al. 2022).”From BIG-bench · §What is in BIG-bench?
- Emergent abilities2022 · cited 6דIn many cases, the effect of scale on performance can often be methodologically predicted via scaling laws—for example, scaling curves for cross-entropy loss have been shown to empirically span more than seven orders of…”From Emergent abilities · §Introduction
- U-PaLM2022 · cited 8דTo this end, large language models not only continue to improve as we scale in terms of data or computational budget (Hoffmann et al. 2022; Kaplan et al. 2020) but also acquire new abilities (Wei et al. 2022a).”From U-PaLM · §Related Work
Abstract
We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant. By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled. We test this hypothesis by training a predicted compute-optimal model, Chinchilla, that uses the same compute budget as Gopher but with 70B parameters and 4$\times$ more more data. Chinchilla uniformly and significantly outperforms Gopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on a large range of downstream evaluation tasks. This also means that Chinchilla uses substantially less compute for fine-tuning and inference, greatly facilitating downstream usage. As a highlight, Chinchilla reaches a state-of-the-art average accuracy of 67.5% on the MMLU benchmark, greater than a 7% improvement over Gopher.