Efficient Large Scale Language Modeling with Mixtures of Experts
Mixture of Experts layers (MoEs) enable efficient scaling of language models through conditional computation. This paper presents a detailed empirical study of how autoregressive MoE language models scale in comparison with dense models in a wide range of settings: in- and out-of-domain language modeling, zero- and few-shot priming, and full-shot fine-tuning.
Also cited · not yet reviewed (17)
- GPT-32020 · cited 25×, 19 in Method“We train autoregressive (decoder-only) transformer models that roughly match the sizes and architecture explored in Brown et al. 2020.”From this paper · §Experimental Setup
- GShard2020 · cited 7×, 4 in Method“Our MoE models follow the design proposed in Lepikhin et al. 2021 with alternating dense and expert layers and top-2 expert selection.”From this paper · §Experimental Setup
- RoBERTa2019 · cited 5×, 4 in Method“We pretrain our models on a union of six English-language datasets, including the five datasets used to pretrain RoBERTa (Liu et al. 2019) and the English subset of CC100, totalling 112B tokens corresponding to 453GB:”From this paper · §Experimental Setup
- BERT2018 · cited 4×, 3 in Method“This is in contrast to the traditional approach of augmenting LMs with task-specific heads, followed by supervised fine-tuning (Devlin et al. 2019; Raffel et al. 2020).”From this paper · §Background and Related Work
Show 13 more
- Switch Transformer2021 · cited 5×, 2 in Method“Recent work (Lewis et al. 2021; Lepikhin et al. 2021; Fedus et al. 2021; Fan et al. 2021) has studied different conditional compute strategies that work well with Transformer models for natural language tasks.”From this paper · §Background and Related Work
- T52019 · cited 4×, 2 in Method“This is in contrast to the traditional approach of augmenting LMs with task-specific heads, followed by supervised fine-tuning (Devlin et al. 2019; Raffel et al. 2020).”From this paper · §Background and Related Work
- The Pile2020 · cited 3×, 2 in Method“For out-of-domain we use data from The Pile (Gao et al. 2021), a public dataset that combines data from 22 diverse sources (e.g., ArXiv, Github, OpenSubtitles, etc.).”From this paper · §Experimental Setup
- BookCorpus (books & movies)2015 · cited 1×, 1 in Method“BookCorpus (Zhu et al. 2019) consists of more than 10K unpublished books (4GB);”From this paper · §Experimental Setup
- Transformer2017 · cited 1×, 1 in Method“While numerous variations have been proposed, such LMs are predominantly based on the transformer architecture (Vaswani et al. 2017).”From this paper · §Background and Related Work
- ReCoRD2018 · cited 1×, 1 in Method“Zero-shot: in addition to the 3 few-shot tasks, we evaluate on ReCoRD (Zhang et al. 2018), HellaSwag (Zellers et al. 2019) and PIQA (Bisk et al. 2020).”From this paper · §Experimental Setup
- Adapters2019 · cited 1×, 1 in Method“Finally, Schick and Schütze 2021b perform few-shot learning by few-shot fine-tuning using pattern-exploiting training, whose efficiency can be improved by performing partial fine-tuning of a small number of additional ta…”From this paper · §Background and Related Work
- BoolQ2019 · cited 1×, 1 in Method“In addition to our zero-shot tasks, we also evaluate on 3 widely-used classification tasks: BoolQ (Clark et al. 2019), MNLI (Williams et al. 2018) and SST-2 (Socher et al. 2013).”From this paper · §Experimental Setup
- BART2019 · cited 1×, 1 in Method“Models are pretrained by hiding parts of the input: predicting the next word sequentially left-to-right, masking words in the text (Devlin et al. 2019; Liu et al. 2019), or perturbing and/or masking spans (Lewis et al. 2…”From this paper · §Background and Related Work
- Prefix-Tuning2021 · cited 1×, 1 in Method“Finally, Schick and Schütze 2021b perform few-shot learning by few-shot fine-tuning using pattern-exploiting training, whose efficiency can be improved by performing partial fine-tuning of a small number of additional ta…”From this paper · §Background and Related Work
- Prompt tuning2021 · cited 1×, 1 in Method“Finally, Schick and Schütze 2021b perform few-shot learning by few-shot fine-tuning using pattern-exploiting training, whose efficiency can be improved by performing partial fine-tuning of a small number of additional ta…”From this paper · §Background and Related Work
- GLaM2021 · cited 1×, 1 in Method“Concurrent to our work, Du et al. 2021, Rajbhandari et al. 2022 and Clark et al. 2022 also study MoE scaling.”From this paper · §Background and Related Work
- Gopher2021 · cited 1×, 1 in Method“Concurrent to our work, Rae et al. 2022 and Smith et al. 2022 further explore scaling dense language models.”From this paper · §Background and Related Work
Led to
- GPT-NeoX-20B2022 · cited 2דWe compare with the GPT-3 models on the OpenAI API (Brown et al. 2020), the open source FairSeq dense models (Artetxe et al. 2021), and GPT-J-6B (Wang and Komatsuzaki 2021).”From GPT-NeoX-20B · §Performance Evaluations
- OPT2022 · cited 3×, 1 in Method“Following Lieber et al. 2021 and Artetxe et al. 2021, we use StereoSet Nadeem et al. 2021 to measure stereotypical bias across 4 categories: profession, gender, religion, and race.”From OPT · §Bias & Toxicity Evaluations
Abstract
Mixture of Experts layers (MoEs) enable efficient scaling of language models through conditional computation. This paper presents a detailed empirical study of how autoregressive MoE language models scale in comparison with dense models in a wide range of settings: in- and out-of-domain language modeling, zero- and few-shot priming, and full-shot fine-tuning. With the exception of fine-tuning, we find MoEs to be substantially more compute efficient. At more modest training budgets, MoEs can match the performance of dense models using $\sim$4 times less compute. This gap narrows at scale, but our largest MoE model (1.1T parameters) consistently outperforms a compute-equivalent dense model (6.7B parameters). Overall, this performance gap varies greatly across tasks and domains, suggesting that MoE and dense models generalize differently in ways that are worthy of future study. We make our code and models publicly available for research use.