Paper Lineage
Esc
MethodSep 2019arXiv 1909.08053cs.CL

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Mohammad Shoeybi, Mostofa Patwary, Raul Puri and 3 others

Recent work in language modeling demonstrates that training large transformer models advances the state of the art in Natural Language Processing applications. However, very large models can be quite difficult to train due to memory constraints.

From the abstract

Built on

12 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (12)

  • BERT2018 · cited 10×, 6 in Method
    “To analyze the effect of model size scaling on accuracy, we train both left-to-right GPT-2 (Radford et al. 2019) language models as well as BERT (Devlin et al. 2018) bidirectional transformers and evaluate them on severa…”
    From this paper · §Introduction
  • ALBERT2019 · cited 9×, 4 in Method
    “For BERT models, we largely follow the training process described in (Lan et al. 2019).”
    From this paper · §Setup
  • Transformer2017 · cited 3×, 3 in Method
    “Current work in NLP trends towards using transformer models (Vaswani et al. 2017) due to their superior accuracy and compute efficiency.”
    From this paper · §Background and Challenges
  • GPipe2018 · cited 4×, 2 in Method
    “It does not require a compiler, and is orthogonal and complementary to the pipeline model parallelism advocated by approaches such as (Huang et al. 2018).”
    From this paper · §Model Parallel Transformers
Show 8 more
  • RoBERTa2019 · cited 3×, 1 in Method
    “For finetuning, we follow the same procedure as (Liu et al. 2019b).”
    From this paper · §Experiments
  • ELMo2018 · cited 2×, 1 in Method
    “By finetuning these pretrained language models on downstream natural language tasks, one can achieve state of the art results as shown in recent work (Devlin et al. 2018; Peters et al. 2018; Howard & Ruder 2018; Radford…”
    From this paper · §Introduction
  • Transformer-XL2019 · cited 2×, 1 in Method
    “Recent parallel work (Ramachandran et al. 2016; Howard & Ruder 2018; Radford et al. 2018; Devlin et al. 2018; Liu et al. 2019b; Dai et al. 2019; Yang et al. 2019; Liu et al. 2019a; Lan et al. 2019) further builds upon th…”
    From this paper · §Background and Challenges
  • XLNet2019 · cited 2×, 1 in Method
    “Recent parallel work (Ramachandran et al. 2016; Howard & Ruder 2018; Radford et al. 2018; Devlin et al. 2018; Liu et al. 2019b; Dai et al. 2019; Yang et al. 2019; Liu et al. 2019a; Lan et al. 2019) further builds upon th…”
    From this paper · §Background and Challenges
  • BookCorpus (books & movies)2015 · cited 1×, 1 in Method
    “For BERT models we include BooksCorpus (Zhu et al. 2015) in the training dataset, however, this dataset is excluded for GPT-2 trainings as it overlaps with LAMBADA task.”
    From this paper · §Setup
  • Goyal large-batch SGD2017 · cited 1×, 1 in Method
    “Further research (Goyal et al. 2017; You et al. 2017; You et al. 2019) has developed techniques to mitigate these effects and drive down the training time of large neural networks.”
    From this paper · §Background and Challenges
  • Learned in Translation2017 · cited 1×, 1 in Method
    “Later work advanced research in this area by learning and transferring neural models that capture contextual representations of words (Melamud et al. 2016; McCann et al. 2017; Peters et al. 2018; Radford et al. 2017; Rad…”
    From this paper · §Background and Challenges
  • AdamW2017 · cited 1×, 1 in Method
    “For our optimizer we utilize Adam (Kingma & Ba 2014) with weight decay (Loshchilov & Hutter 2019) λ=0.01\lambda=0.01.”
    From this paper · §Setup

Led to

  • GPT-32020 · cited 3×
    “Work in this vein has successively increased model size: 213 million parameters [134] in the original paper, 300 million parameters [20], 1.5 billion parameters [117], 8 billion parameters [125], 11 billion parameters [1…”
    From GPT-3 · §Related Work
  • VirTex2020 · cited 2×, 1 in Method
    “However, following BERT [64], many large-scale models [82, 83] instead use masked language models (MLMs): some tokens are randomly masked and are predicted by the model.”
    From VirTex · §Method
  • The Pile2020 · cited 2×
    “More recently, language models (Radford et al. 2018; Radford et al. 2019; Brown et al. 2020; Rosset 2019; Shoeybi et al. 2019) and masked language models (Devlin et al. 2019; Liu et al. 2019; Raffel et al. 2019) have bee…”
    From The Pile · §Related Work
  • Gopher2021 · cited 2×, 2 in Method
    “To address these memory concerns, we use optimiser state partitioning (Rajbhandari et al. 2020), model parallelism (Shoeybi et al. 2019), and rematerialisation (Griewank and Walther 2000) to partition the model state and…”
    From Gopher · §Method
  • Megatron-Turing NLG2022 · cited 5×, 3 in Method
    “Megatron [63] uses model parallelism to efficiently partition transformer blocks for large-scale language models.”
    From Megatron-Turing NLG · §Large Model Training Infrastructure
  • GPT-NeoX-20B2022 · cited 3×, 2 in Method
    “Our model is trained using a codebase that builds on Megatron (Shoeybi et al. 2020) and DeepSpeed (Rasley et al. 2020) to facilitate efficient and straightforward training of large language models with tens of billions o…”
    From GPT-NeoX-20B · §Model Design and Implementation
  • OPT2022 · cited 2×, 1 in Method
    “We trained OPT-175B on 992 80GB A100 GPUs, by utilizing Fully Sharded Data Parallel Artetxe et al. 2021 with Megatron-LM Tensor Parallelism Shoeybi et al. 2019.”
    From OPT · §Method
Abstract

Recent work in language modeling demonstrates that training large transformer models advances the state of the art in Natural Language Processing applications. However, very large models can be quite difficult to train due to memory constraints. In this work, we present our techniques for training very large transformer models and implement a simple, efficient intra-layer model parallel approach that enables training transformer models with billions of parameters. Our approach does not require a new compiler or library changes, is orthogonal and complimentary to pipeline model parallelism, and can be fully implemented with the insertion of a few communication operations in native PyTorch. We illustrate this approach by converging transformer based models up to 8.3 billion parameters using 512 GPUs. We sustain 15.1 PetaFLOPs across the entire application with 76% scaling efficiency when compared to a strong single GPU baseline that sustains 39 TeraFLOPs, which is 30% of peak FLOPs. To demonstrate that large language models can further advance the state of the art (SOTA), we train an 8.3 billion parameter transformer language model similar to GPT-2 and a 3.9 billion parameter model similar to BERT. We show that careful attention to the placement of layer normalization in BERT-like models is critical to achieving increased performance as the model size grows. Using the GPT-2 model we achieve SOTA results on the WikiText103 (10.8 compared to SOTA perplexity of 15.8) and LAMBADA (66.5% compared to SOTA accuracy of 63.2%) datasets. Our BERT model achieves SOTA results on the RACE dataset (90.9% compared to SOTA accuracy of 89.4%).