Paper Lineage
Esc
MethodDec 2021arXiv 2112.06905cs.CL

GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Nan Du, Yanping Huang, Andrew M. Dai and 24 others

Scaling language models with more data, compute and parameters has driven significant progress in natural language processing. For example, thanks to scaling, GPT-3 was able to achieve strong results on in-context learning tasks.

From the abstract

Built on

10 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (10)

  • GPT-32020 · cited 13×, 3 in Method
    “GPT-3 (Brown et al. 2020) and related work (Shoeybi et al. 2019; Lieber et al. 2021; Wei et al. 2021) demonstrated that scaling up language models greatly improves task-agnostic, few-shot performance.”
    From this paper · §Related Work
  • GShard2020 · cited 4×, 2 in Method
    “Similar to the GShard MoE Transformer (Lepikhin et al. 2021), we replace the feed-forward component of every other Transformer layer with an MoE layer, as shown in Figure 2.”
    From this paper · §Model Architecture
  • Switch Transformer2021 · cited 4×, 1 in Method
    “We leverage sparsely activated Mixture-of-Experts (MoE) (Shazeer et al. 2017; Fedus et al. 2021) in GLaM models.”
    From this paper · §Model Architecture
  • SentencePiece2018 · cited 1×, 1 in Method
    “We use the SentencePiece (Kudo & Richardson 2018) subword tokenizer with a vocabulary of size of 256256K.”
    From this paper · §Experiment Setup
Show 6 more
  • ReCoRD2018 · cited 1×, 1 in Method
    “On a few tasks, such as ReCoRD (Zhang et al. 2018) and COPA (Gordon et al. 2012), the non-normalized loss can yield better results and thus is adopted.”
    From this paper · §Experiment Setup
  • Transformer-XL2019 · cited 1×, 1 in Method
    “We replace the standard positional embedding with per-layer relative positional bias from Dai et al. 2019.”
    From this paper · §Model Architecture
  • word2vec2013 · cited 2×
    “Word embedding models and extensions such as word2vec (Mikolov et al. 2013), GloVe (Pennington et al. 2014) and paragraph vectors (Le & Mikolov 2014) have shown good generalization to many tasks simply by transferring th…”
    From this paper · §Related Work
  • BERT2018 · cited 2×
    “More recently, models that used Transformers (Vaswani et al. 2017) showed that larger models with self-supervision on unlabeled data could yield significant improvements on NLP tasks (Devlin et al. 2019; Yang et al. 2019…”
    From this paper · §Related Work
  • FLAN2021 · cited 2×
    “GPT-3 (Brown et al. 2020) and related work (Shoeybi et al. 2019; Lieber et al. 2021; Wei et al. 2021) demonstrated that scaling up language models greatly improves task-agnostic, few-shot performance.”
    From this paper · §Related Work
  • Gopher2021 · cited 2×
    “We also analyse toxicity degeneration similarly to Gopher (Rae et al. 2021), and extend the analysis to consider the human-behavioral baseline.”
    From this paper · §Ethics and Unintended Biases

Led to

  • Fairseq MoE LMs2021 · cited 1×, 1 in Method
    “Concurrent to our work, Du et al. 2021, Rajbhandari et al. 2022 and Clark et al. 2022 also study MoE scaling.”
    From Fairseq MoE LMs · §Background and Related Work
  • Chinchilla2022 · cited 2×
    “These include both dense transformer models (Brown et al. 2020; Lieber et al. 2021; Smith et al. 2022; Rae et al. 2021; Thoppilan et al. 2022) and mixture-of-expert (MoE) models (Du et al. 2021; Fedus et al. 2021; Zoph e…”
    From Chinchilla · §Related Work
  • PaLM2022 · cited 12×
    “The most powerful of these post-GPT-3 models are GLaM (Du et al. 2021), Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), Megatron–Turing NLG (Smith et al. 2022), and LaMDA (Thoppilan et al. 2022), all of whic…”
    From PaLM · §Introduction
  • Emergent abilities2022 · cited 2×
    “For example, Chinchilla (Hoffmann et al. 2022) has one-fourth as many parameters as Gopher (Rae et al. 2021) but uses similar training compute; and sparse mixture-of-expert models have more parameters per training/infere…”
    From Emergent abilities · §Emergent Abilities Definition
  • Flan-T5 / Flan-PaLM2022 · cited 2×
    “These benchmarks were also used in the PaLM paper (Chowdhery et al. 2022), which did not find any meaningful data contamination with pre-training data, consistent with data contamination analyses in previous work (Brown…”
    From Flan-T5 / Flan-PaLM · §unknown section
Abstract

Scaling language models with more data, compute and parameters has driven significant progress in natural language processing. For example, thanks to scaling, GPT-3 was able to achieve strong results on in-context learning tasks. However, training these large dense models requires significant amounts of computing resources. In this paper, we propose and develop a family of language models named GLaM (Generalist Language Model), which uses a sparsely activated mixture-of-experts architecture to scale the model capacity while also incurring substantially less training cost compared to dense variants. The largest GLaM has 1.2 trillion parameters, which is approximately 7x larger than GPT-3. It consumes only 1/3 of the energy used to train GPT-3 and requires half of the computation flops for inference, while still achieving better overall zero-shot and one-shot performance across 29 NLP tasks.