Transcending Scaling Laws with 0.1% Extra Compute
Scaling language models improves performance but comes with significant computational costs. This paper proposes UL2R, a method that substantially improves existing language models and their scaling curves with a relatively tiny amount of extra compute.
Also cited · not yet reviewed (13)
- Chinchilla2022 · cited 8דTo this end, large language models not only continue to improve as we scale in terms of data or computational budget (Hoffmann et al. 2022; Kaplan et al. 2020) but also acquire new abilities (Wei et al. 2022a).”From this paper · §Related Work
- Emergent abilities2022 · cited 7דTo this end, large language models not only continue to improve as we scale in terms of data or computational budget (Hoffmann et al. 2022; Kaplan et al. 2020) but also acquire new abilities (Wei et al. 2022a).”From this paper · §Related Work
- T52019 · cited 6דThe majority of this prior work, however, requires additional data such as aggregating dozens or hundreds of NLP datasets (Raffel et al. 2019; Aghajanyan et al. 2021; Aribandi et al. 2022), writing additional templates o…”From this paper · §Related Work
- PaLM2022 · cited 6דScaling and improving large language models is one of the most impactful research areas in modern artificial intelligence (Chowdhery et al. 2022).”From this paper · §Related Work
Show 9 more
- Kaplan scaling laws2020 · cited 5דTo this end, large language models not only continue to improve as we scale in terms of data or computational budget (Hoffmann et al. 2022; Kaplan et al. 2020) but also acquire new abilities (Wei et al. 2022a).”From this paper · §Related Work
- GPT-32020 · cited 4דFor evaluation, we use the average score of NLU and NLG tasks from the GPT-3 suite (Brown et al. 2020).”From this paper · §Experiments
- Gopher2021 · cited 3דAside from PaLM 62B and PaLM 540B which we use for direct comparisons with U-PaLM, we also compare with Chinchilla 70B (Hoffmann et al. 2022) and Gopher 280B (Rae et al. 2021).”From this paper · §Experiments
- BERT2018 · cited 2דWhile there have been many paradigms and self-supervision methods proposed to train these models (Devlin et al. 2018; Clark et al. 2020b; Yang et al. 2019; Raffel et al. 2019), to this date most large language models (i.…”From this paper · §Related Work
- FLAN2021 · cited 2דA range of prior work has shown that finetuning language models on a collection of NLP tasks can improve downstream performance on a broad range of downstream tasks (Aghajanyan et al. 2021; Aribandi et al. 2022; Wei et a…”From this paper · §Related Work
- T02021 · cited 2דA range of prior work has shown that finetuning language models on a collection of NLP tasks can improve downstream performance on a broad range of downstream tasks (Aghajanyan et al. 2021; Aribandi et al. 2022; Wei et a…”From this paper · §Related Work
- GSM8K verifiers2021 · cited 2דWe use the GSM8K (Cobbe et al. 2021), BBH (Suzgun et al. 2022), StrategyQA (Geva et al. 2021) and CommonsenseQA (Talmor et al. 2019) benchmarks.”From this paper · §Experiments
- InstructGPT2022 · cited 2דA range of prior work has shown that finetuning language models on a collection of NLP tasks can improve downstream performance on a broad range of downstream tasks (Aghajanyan et al. 2021; Aribandi et al. 2022; Wei et a…”From this paper · §Related Work
- Flan-T5 / Flan-PaLM2022 · cited 2דNormally, we would like to end with a cliche future work statement here but not today because we have done this already here (Chung et al. 2022).”From this paper · §Conclusion and Future Work
Led to
- Flan-T5 / Flan-PaLM2022 · cited 4דFinally, we instruction-finetune U-PaLM, which is a 540B PaLM model initialized from PaLM-540B and then pretrained with an UL2 objective for 20k additional steps (Tay et al. 2022a; Tay et al. 2022b).”From Flan-T5 / Flan-PaLM · §unknown section
Abstract
Scaling language models improves performance but comes with significant computational costs. This paper proposes UL2R, a method that substantially improves existing language models and their scaling curves with a relatively tiny amount of extra compute. The key idea is to continue training a state-of-the-art large language model (e.g., PaLM) on a few more steps with UL2's mixture-of-denoiser objective. We show that, with almost negligible extra computational costs and no new sources of data, we are able to substantially improve the scaling properties of large language models on downstream metrics. In this paper, we continue training PaLM with UL2R, introducing a new set of models at 8B, 62B, and 540B scale which we call U-PaLM. Impressively, at 540B scale, we show an approximately 2x computational savings rate where U-PaLM achieves the same performance as the final PaLM 540B model at around half its computational budget (i.e., saving $\sim$4.4 million TPUv4 hours). We further show that this improved scaling curve leads to 'emergent abilities' on challenging BIG-Bench tasks -- for instance, U-PaLM does much better than PaLM on some tasks or demonstrates better quality at much smaller scale (62B as opposed to 540B). Overall, we show that U-PaLM outperforms PaLM on many few-shot setups, i.e., English NLP tasks (e.g., commonsense reasoning, question answering), reasoning tasks with chain-of-thought (e.g., GSM8K), multilingual tasks (MGSM, TydiQA), MMLU and challenging BIG-Bench tasks. Finally, we provide qualitative examples showing the new capabilities of U-PaLM for single and multi-span infilling.