Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world. In this paper, we present an analysis of Transformer-based language model performance across a wide range of model scales -- from models with tens of millions of parameters up to a 280 billion parameter model called Gopher.
Also cited · not yet reviewed (14)
- BERT2018 · cited 2×, 2 in Method“Whilst there are other objectives towards modelling a sequence, such as modelling masked tokens given bi-directional context (Mikolov et al. 2013; Devlin et al. 2019) and modelling all permutations of the sequence (Yang…”From this paper · §Background
- Transformer-XL2019 · cited 2×, 2 in Method“We use the autoregressive Transformer architecture detailed in Radford et al. 2019 with two modifications: we use RMSNorm (Zhang and Sennrich 2019) instead of LayerNorm (Ba et al. 2016), and we use the relative positiona…”From this paper · §Method
- Megatron-LM2019 · cited 2×, 2 in Method“To address these memory concerns, we use optimiser state partitioning (Rajbhandari et al. 2020), model parallelism (Shoeybi et al. 2019), and rematerialisation (Griewank and Walther 2000) to partition the model state and…”From this paper · §Method
- T52019 · cited 2×, 2 in Method“The largest sampling subset is our curated web-text corpus MassiveWeb, which we find to improve downstream performance relative to existing web-text datasets such as C4 (Raffel et al. 2020b) in Figure A5.”From this paper · §Method
Show 10 more
- Switch Transformer2021 · cited 3×, 1 in Method“There have also been Transformer variants which incorporate a sparse mixture of experts (Fedus et al. 2021; Roller et al. 2021b) to increase the model size (in some cases to trillions of parameters) with more modest comp…”From this paper · §Background
- Transformer2017 · cited 2×, 1 in Method“Progress has been driven by both scale and network architecture (Hochreiter and Schmidhuber 1997; Bahdanau et al. 2014; Vaswani et al. 2017). rosenfeld2020a and Kaplan et al. 2020 independently found power laws relating…”From this paper · §Introduction
- word2vec2013 · cited 1×, 1 in Method“Whilst there are other objectives towards modelling a sequence, such as modelling masked tokens given bi-directional context (Mikolov et al. 2013; Devlin et al. 2019) and modelling all permutations of the sequence (Yang…”From this paper · §Background
- SentencePiece2018 · cited 1×, 1 in Method“We tokenize the text using SentencePiece (Kudo and Richardson 2018) with a vocabulary of 32,000 and use a byte-level backoff to support open-vocabulary modelling.”From this paper · §Method
- GPipe2018 · cited 1×, 1 in Method“Therefore, we find that pipelining (Huang et al. 2019) is not necessary on TPUs until the training scale exceeds the 1024-chip “pod”, which greatly simplifies training mid-sized models.”From this paper · §Method
- XLNet2019 · cited 1×, 1 in Method“Whilst there are other objectives towards modelling a sequence, such as modelling masked tokens given bi-directional context (Mikolov et al. 2013; Devlin et al. 2019) and modelling all permutations of the sequence (Yang…”From this paper · §Background
- The Pile2020 · cited 1×, 1 in Method“Since GPT-3 there has been a 178B parameter Transformer language model Jurassic-1 (Lieber et al. 2021) which uses a diverse training set and a larger tokenizer vocabulary size, along with an announced 530B Megatron-Turin…”From this paper · §Background
- FLAN2021 · cited 1×, 1 in Method“Other recent LLMs include two models (FLAN and T0) fine-tuned on instructions for an array of down-stream tasks (Sanh et al. 2021; Wei et al. 2021) which improves performance to unseen tasks — these ideas are complementa…”From this paper · §Background
- T02021 · cited 1×, 1 in Method“Other recent LLMs include two models (FLAN and T0) fine-tuned on instructions for an array of down-stream tasks (Sanh et al. 2021; Wei et al. 2021) which improves performance to unseen tasks — these ideas are complementa…”From this paper · §Background
- GPT-32020 · cited 7דThe baseline model includes LLMs such as GPT-3 (175B parameters) (Brown et al. 2020), Jurassic-1 (Lieber et al. 2021) (178B parameters), and Megatron-Turing NLG (530B parameters) (Kharya and Alvi 2021); the exact baselin…”From this paper · §Results
Led to
- GLaM2021 · cited 2דWe also analyse toxicity degeneration similarly to Gopher (Rae et al. 2021), and extend the analysis to consider the human-behavioral baseline.”From GLaM · §Ethics and Unintended Biases
- Fairseq MoE LMs2021 · cited 1×, 1 in Method“Concurrent to our work, Rae et al. 2022 and Smith et al. 2022 further explore scaling dense language models.”From Fairseq MoE LMs · §Background and Related Work
- Megatron-Turing NLG2022 · cited 5דTo evaluate the quality of our model (as well as other pretrained language models), we adopt a zero-/one-/few-shot evaluation setting similar to prior work [9, 53].”From Megatron-Turing NLG · §Results and Achievements
- InstructGPT2022 · cited 2×, 1 in Method“This is because the language modeling objective used for many recent large LMs—predicting the next token on a webpage from the internet—is different from the objective “follow the user’s instructions helpfully and safely…”From InstructGPT · §Introduction
- Chinchilla2022 · cited 16דThese include both dense transformer models (Brown et al. 2020; Lieber et al. 2021; Smith et al. 2022; Rae et al. 2021; Thoppilan et al. 2022) and mixture-of-expert (MoE) models (Du et al. 2021; Fedus et al. 2021; Zoph e…”From Chinchilla · §Related Work
- PaLM2022 · cited 11דThe most powerful of these post-GPT-3 models are GLaM (Du et al. 2021), Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), Megatron–Turing NLG (Smith et al. 2022), and LaMDA (Thoppilan et al. 2022), all of whic…”From PaLM · §Introduction
- OPT2022 · cited 9דLarge language models (LLMs) trained on massive text collections have shown surprising emergent capabilities to generate text and perform zero- and few-shot learning Brown et al. 2020; Lieber et al. 2021; Smith et al. 20…”From OPT · §Introduction
- BIG-bench2022 · cited 2דBIG-bench has also been partially evaluated on other models including Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), and T0 (Sanh et al. 2022).”From BIG-bench · §What is in BIG-bench?
- Emergent abilities2022 · cited 5דTraining dataset size is also an important factor, but we do not plot capabilities against it because many language model families use a fixed number of training examples for all model sizes (Brown et al. 2020; Rae et al…”From Emergent abilities · §Emergent Abilities Definition
- U-PaLM2022 · cited 3דAside from PaLM 62B and PaLM 540B which we use for direct comparisons with U-PaLM, we also compare with Chinchilla 70B (Hoffmann et al. 2022) and Gopher 280B (Rae et al. 2021).”From U-PaLM · §Experiments
Abstract
Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world. In this paper, we present an analysis of Transformer-based language model performance across a wide range of model scales -- from models with tens of millions of parameters up to a 280 billion parameter model called Gopher. These models are evaluated on 152 diverse tasks, achieving state-of-the-art performance across the majority. Gains from scale are largest in areas such as reading comprehension, fact-checking, and the identification of toxic language, but logical and mathematical reasoning see less benefit. We provide a holistic analysis of the training dataset and model's behaviour, covering the intersection of model scale with bias and toxicity. Finally we discuss the application of language models to AI safety and the mitigation of downstream harms.