GLaM: Efficient Scaling of Language Models with Mixture-of-Experts
Scaling language models with more data, compute and parameters has driven significant progress in natural language processing. For example, thanks to scaling, GPT-3 was able to achieve strong results on in-context learning tasks.
Also cited · not yet reviewed (10)
- GPT-32020 · cited 13×, 3 in Method“GPT-3 (Brown et al. 2020) and related work (Shoeybi et al. 2019; Lieber et al. 2021; Wei et al. 2021) demonstrated that scaling up language models greatly improves task-agnostic, few-shot performance.”From this paper · §Related Work
- GShard2020 · cited 4×, 2 in Method“Similar to the GShard MoE Transformer (Lepikhin et al. 2021), we replace the feed-forward component of every other Transformer layer with an MoE layer, as shown in Figure 2.”From this paper · §Model Architecture
- Switch Transformer2021 · cited 4×, 1 in Method“We leverage sparsely activated Mixture-of-Experts (MoE) (Shazeer et al. 2017; Fedus et al. 2021) in GLaM models.”From this paper · §Model Architecture
- SentencePiece2018 · cited 1×, 1 in Method“We use the SentencePiece (Kudo & Richardson 2018) subword tokenizer with a vocabulary of size of 256256K.”From this paper · §Experiment Setup
Show 6 more
- ReCoRD2018 · cited 1×, 1 in Method“On a few tasks, such as ReCoRD (Zhang et al. 2018) and COPA (Gordon et al. 2012), the non-normalized loss can yield better results and thus is adopted.”From this paper · §Experiment Setup
- Transformer-XL2019 · cited 1×, 1 in Method“We replace the standard positional embedding with per-layer relative positional bias from Dai et al. 2019.”From this paper · §Model Architecture
- word2vec2013 · cited 2דWord embedding models and extensions such as word2vec (Mikolov et al. 2013), GloVe (Pennington et al. 2014) and paragraph vectors (Le & Mikolov 2014) have shown good generalization to many tasks simply by transferring th…”From this paper · §Related Work
- BERT2018 · cited 2דMore recently, models that used Transformers (Vaswani et al. 2017) showed that larger models with self-supervision on unlabeled data could yield significant improvements on NLP tasks (Devlin et al. 2019; Yang et al. 2019…”From this paper · §Related Work
- FLAN2021 · cited 2דGPT-3 (Brown et al. 2020) and related work (Shoeybi et al. 2019; Lieber et al. 2021; Wei et al. 2021) demonstrated that scaling up language models greatly improves task-agnostic, few-shot performance.”From this paper · §Related Work
- Gopher2021 · cited 2דWe also analyse toxicity degeneration similarly to Gopher (Rae et al. 2021), and extend the analysis to consider the human-behavioral baseline.”From this paper · §Ethics and Unintended Biases
Led to
- Fairseq MoE LMs2021 · cited 1×, 1 in Method“Concurrent to our work, Du et al. 2021, Rajbhandari et al. 2022 and Clark et al. 2022 also study MoE scaling.”From Fairseq MoE LMs · §Background and Related Work
- Chinchilla2022 · cited 2דThese include both dense transformer models (Brown et al. 2020; Lieber et al. 2021; Smith et al. 2022; Rae et al. 2021; Thoppilan et al. 2022) and mixture-of-expert (MoE) models (Du et al. 2021; Fedus et al. 2021; Zoph e…”From Chinchilla · §Related Work
- PaLM2022 · cited 12דThe most powerful of these post-GPT-3 models are GLaM (Du et al. 2021), Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), Megatron–Turing NLG (Smith et al. 2022), and LaMDA (Thoppilan et al. 2022), all of whic…”From PaLM · §Introduction
- Emergent abilities2022 · cited 2דFor example, Chinchilla (Hoffmann et al. 2022) has one-fourth as many parameters as Gopher (Rae et al. 2021) but uses similar training compute; and sparse mixture-of-expert models have more parameters per training/infere…”From Emergent abilities · §Emergent Abilities Definition
- Flan-T5 / Flan-PaLM2022 · cited 2דThese benchmarks were also used in the PaLM paper (Chowdhery et al. 2022), which did not find any meaningful data contamination with pre-training data, consistent with data contamination analyses in previous work (Brown…”From Flan-T5 / Flan-PaLM · §unknown section
Abstract
Scaling language models with more data, compute and parameters has driven significant progress in natural language processing. For example, thanks to scaling, GPT-3 was able to achieve strong results on in-context learning tasks. However, training these large dense models requires significant amounts of computing resources. In this paper, we propose and develop a family of language models named GLaM (Generalist Language Model), which uses a sparsely activated mixture-of-experts architecture to scale the model capacity while also incurring substantially less training cost compared to dense variants. The largest GLaM has 1.2 trillion parameters, which is approximately 7x larger than GPT-3. It consumes only 1/3 of the energy used to train GPT-3 and requires half of the computation flops for inference, while still achieving better overall zero-shot and one-shot performance across 29 NLP tasks.