Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
In deep learning, models typically reuse the same parameters for all inputs. Mixture of Experts (MoE) defies this and instead selects different parameters for each incoming example.
Also cited · not yet reviewed (8)
- T52019 · cited 12×, 2 in Method“An approach followed in Radford et al. 2018; Raffel et al. 2019; Brown et al. 2020 expands the model size of a densely-activated Transformer (Vaswani et al. 2017).”From this paper · §Introduction
- GPT-32020 · cited 4×, 1 in Method“It is common to mix both model and data parallelism for large scale models, which was done in the largest T5 models (Raffel et al. 2019; Xue et al. 2020) and in GPT-3 (Brown et al. 2020).”From this paper · §Designing Models with Data, Model, and Expert-Parallelism
- XLNet2019 · cited 1×, 1 in Method“On ANLI (Nie et al. 2019), Switch XXL improves over the prior state-of-the-art to get a 65.7 accuracy versus the prior best of 49.4 (Yang et al. 2020).”From this paper · §Designing Models with Data, Model, and Expert-Parallelism
- Kaplan scaling laws2020 · cited 6דLarge scale training has been an effective path towards flexible and powerful neural language models (Radford et al. 2018; Kaplan et al. 2020; Brown et al. 2020).”From this paper · §Introduction
Show 4 more
- GShard2020 · cited 5דMoE models have had notable successes in machine translation (Shazeer et al. 2017; Shazeer et al. 2018; Lepikhin et al. 2020), however, widespread adoption is hindered by complexity, communication costs, and training ins…”From this paper · §Introduction
- Knowledge Distillation2015 · cited 2דFurther, our large sparse models can be distilled (Hinton et al. 2015) into small dense versions while preserving 30% of the sparse model quality gain.”From this paper · §Introduction
- Transformer2017 · cited 2דThe guiding design principle for Switch Transformers is to maximize the parameter count of a Transformer model (Vaswani et al. 2017) in a simple and computationally efficient way.”From this paper · §Switch Transformer
- Knowledge in LM parameters2020 · cited 2דAnd as in Roberts et al. 2020, we evaluate the knowledge of our models by fine-tuning on three closed-book question answering data sets: Natural Questions (Kwiatkowski et al. 2019), Web Questions (Berant et al. 2013) and…”From this paper · §Downstream Results
Led to
- VLMo2021 · cited 1×, 1 in Method“Inspired by mixture-of-experts networks [40, 13], we propose a general-purpose multimodal Transformer for vision-language tasks, namely MoME Transformer, to encode different modalities.”From VLMo · §Methods
- Swin V22021 · cited 3דIt significantly improves a model’s performance on language tasks devlin2018bert; radford2019language; raffel2019t5; Turing-17B; fedus2021switch; Megatron-Turing-530B and the model demonstrates amazing few-shot capabilit…”From Swin V2 · §Introduction
- Gopher2021 · cited 3×, 1 in Method“There have also been Transformer variants which incorporate a sparse mixture of experts (Fedus et al. 2021; Roller et al. 2021b) to increase the model size (in some cases to trillions of parameters) with more modest comp…”From Gopher · §Background
- GLaM2021 · cited 4×, 1 in Method“We leverage sparsely activated Mixture-of-Experts (MoE) (Shazeer et al. 2017; Fedus et al. 2021) in GLaM models.”From GLaM · §Model Architecture
- Fairseq MoE LMs2021 · cited 5×, 2 in Method“Recent work (Lewis et al. 2021; Lepikhin et al. 2021; Fedus et al. 2021; Fan et al. 2021) has studied different conditional compute strategies that work well with Transformer models for natural language tasks.”From Fairseq MoE LMs · §Background and Related Work
- InstructGPT2022 · cited 2×, 1 in Method“This is because the language modeling objective used for many recent large LMs—predicting the next token on a webpage from the internet—is different from the objective “follow the user’s instructions helpfully and safely…”From InstructGPT · §Introduction
- Chinchilla2022 · cited 2דThese include both dense transformer models (Brown et al. 2020; Lieber et al. 2021; Smith et al. 2022; Rae et al. 2021; Thoppilan et al. 2022) and mixture-of-expert (MoE) models (Du et al. 2021; Fedus et al. 2021; Zoph e…”From Chinchilla · §Related Work
- PaLM2022 · cited 2דScale has continued to increase after GPT-3, evidenced by the succession of the 178B parameter Jurassic-1 (Lieber et al. 2021), the 280B parameter Gopher model (Rae et al. 2021), the 530B Megatron-Turing NLG (Smith et al…”From PaLM · §Related Work
- BIG-bench2022 · cited 2דMotivated by this predictable improvement, researchers have now scaled language models to more than one trillion parameters (Fedus et al. 2021), and we expect models to grow orders of magnitude larger over the next sever…”From BIG-bench · §Introduction
- Emergent abilities2022 · cited 2דFor example, Chinchilla (Hoffmann et al. 2022) has one-fourth as many parameters as Gopher (Rae et al. 2021) but uses similar training compute; and sparse mixture-of-expert models have more parameters per training/infere…”From Emergent abilities · §Emergent Abilities Definition
Abstract
In deep learning, models typically reuse the same parameters for all inputs. Mixture of Experts (MoE) defies this and instead selects different parameters for each incoming example. The result is a sparsely-activated model -- with outrageous numbers of parameters -- but a constant computational cost. However, despite several notable successes of MoE, widespread adoption has been hindered by complexity, communication costs and training instability -- we address these with the Switch Transformer. We simplify the MoE routing algorithm and design intuitive improved models with reduced communication and computational costs. Our proposed training techniques help wrangle the instabilities and we show large sparse models may be trained, for the first time, with lower precision (bfloat16) formats. We design models based off T5-Base and T5-Large to obtain up to 7x increases in pre-training speed with the same computational resources. These improvements extend into multilingual settings where we measure gains over the mT5-Base version across all 101 languages. Finally, we advance the current scale of language models by pre-training up to trillion parameter models on the "Colossal Clean Crawled Corpus" and achieve a 4x speedup over the T5-XXL model.