Paper Lineage
Esc
MethodJan 2022arXiv 2201.11990cs.CL

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Shaden Smith, Mostofa Patwary, Brandon Norick and 17 others

Pretrained general-purpose language models can achieve state-of-the-art accuracies in various natural language processing domains by adapting to downstream tasks via zero-shot, few-shot and fine-tuning techniques. Because of their success, the size of these models has increased rapidly, requiring high-performance hardware, software, and algorithmic techniques to enable training such large models.

From the abstract

Built on

8 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (8)

  • GPT-32020 · cited 23×, 7 in Method
    “As prior work has found (e.g., [9]), the quality of unfiltered Common Crawl data is lower than that of curated datasets and steps should be taken to increase the average quality of data selected from Common Crawl for LM…”
    From this paper · §Training Dataset and Model Configuration
  • The Pile2020 · cited 5×, 5 in Method
    “To compile our training dataset, we made use of recent work aimed at collecting a diverse training set for language modeling [17].”
    From this paper · §Training Dataset and Model Configuration
  • Megatron-LM2019 · cited 5×, 3 in Method
    “Megatron [63] uses model parallelism to efficiently partition transformer blocks for large-scale language models.”
    From this paper · §Large Model Training Infrastructure
  • GPipe2018 · cited 1×, 1 in Method
    “Pipeline model parallelism (or, pipeline parallelism) divides the layers of the model into stages that can be processed in parallel [23, 42].”
    From this paper · §Large Model Training Infrastructure
Show 4 more
  • Gopher2021 · cited 5×
    “To evaluate the quality of our model (as well as other pretrained language models), we adopt a zero-/one-/few-shot evaluation setting similar to prior work [9, 53].”
    From this paper · §Results and Achievements
  • BERT2018 · cited 2×
    “This trend is continued when large-scale pretraining with transformer architectures becomes popular, with BERT [12] scaling up to 300 million parameters, followed by GPT-2 [52] at 1.5 billion parameters.”
    From this paper · §Related Works
  • T52019 · cited 2×
    “Prior work after GPT-2 produced dense transformer models at 8 billion [63], 11 billion [54], and 17 billion [4] parameters, and GPT-3 [9] at 175 billion parameters demonstrated for the first time that language models at…”
    From this paper · §Related Works
  • FLAN2021 · cited 2×
    “Both T0 [59] and FLAN [70] have taken this path and have shown that such an approach can improve zero-shot learning capabilities of language models.”
    From this paper · §Related Works

Led to

  • Chinchilla2022 · cited 3×
    “These include both dense transformer models (Brown et al. 2020; Lieber et al. 2021; Smith et al. 2022; Rae et al. 2021; Thoppilan et al. 2022) and mixture-of-expert (MoE) models (Du et al. 2021; Fedus et al. 2021; Zoph e…”
    From Chinchilla · §Related Work
  • PaLM2022 · cited 8×
    “The most powerful of these post-GPT-3 models are GLaM (Du et al. 2021), Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), Megatron–Turing NLG (Smith et al. 2022), and LaMDA (Thoppilan et al. 2022), all of whic…”
    From PaLM · §Introduction
  • GPT-NeoX-20B2022 · cited 2×
    “A consequence of this has been an abundance of research focusing on scaling Transformer models up to ever-larger scales, resulting in dense models that surpass 500B parameters (Smith et al. 2022; Chowdhery et al. 2022),…”
    From GPT-NeoX-20B · §Introduction
  • OPT2022 · cited 4×
    “Auto-regressive language models Mikolov et al. 2009 have seen the largest growth in model size, from 117M parameters Radford et al. 2018 to over 500B parameters Smith et al. 2022; Chowdhery et al. 2022.”
    From OPT · §Related Work
Abstract

Pretrained general-purpose language models can achieve state-of-the-art accuracies in various natural language processing domains by adapting to downstream tasks via zero-shot, few-shot and fine-tuning techniques. Because of their success, the size of these models has increased rapidly, requiring high-performance hardware, software, and algorithmic techniques to enable training such large models. As the result of a joint effort between Microsoft and NVIDIA, we present details on the training of the largest monolithic transformer based language model, Megatron-Turing NLG 530B (MT-NLG), with 530 billion parameters. In this paper, we first focus on the infrastructure as well as the 3D parallelism methodology used to train this model using DeepSpeed and Megatron. Next, we detail the training process, the design of our training corpus, and our data curation techniques, which we believe is a key ingredient to the success of the model. Finally, we discuss various evaluation results, as well as other interesting observations and new properties exhibited by MT-NLG. We demonstrate that MT-NLG achieves superior zero-, one-, and few-shot learning accuracies on several NLP benchmarks and establishes new state-of-the-art results. We believe that our contributions will help further the development of large-scale training infrastructures, large-scale language models, and natural language generations.