Paper Lineage
Esc
MethodNov 2018arXiv 1811.06965cs.CV

GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism

Yanping Huang, Youlong Cheng, Ankur Bapna and 8 others

Scaling up deep neural network capacity has been known as an effective approach to improving model quality for several different machine learning tasks. In many cases, increasing model capacity beyond the memory limit of a single accelerator has required developing special algorithms or infrastructure.

From the abstract

Built on

2 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (2)

  • Transformer2017 · cited 4×
    “Our comparison is based on the performance of a single Transformer [15] trained on all language pairs in this corpus.”
    From this paper · §Massive Massively Multilingual Machine Translation
  • BERT2018 · cited 2×
    “A similar phenomenon can also be observed in the context of natural language processing (Figure 1b) where simple shallow models of sentence representations [1, 2] are outperformed by their deeper and larger counterparts…”
    From this paper · §Introduction

Led to

  • Megatron-LM2019 · cited 4×, 2 in Method
    “It does not require a compiler, and is orthogonal and complementary to the pipeline model parallelism advocated by approaches such as (Huang et al. 2018).”
    From Megatron-LM · §Model Parallel Transformers
  • T52019 · cited 2×
    “This is a natural fit for neural networks, which have been shown to exhibit remarkable scalability, i.e. it is often possible to achieve better performance simply by training a larger model on a larger data set (Hestness…”
    From T5 · §Introduction
  • GShard2020 · cited 15×
    “Similarly in natural language processing, scaling Transformers [10] yielded consistent gains on language understanding tasks [4, 11, 12], cross-lingual down-stream transfer [4, 13] and (massively-)multilingual neural mac…”
    From GShard · §Introduction
  • Gopher2021 · cited 1×, 1 in Method
    “Therefore, we find that pipelining (Huang et al. 2019) is not necessary on TPUs until the training scale exceeds the 1024-chip “pod”, which greatly simplifies training mid-sized models.”
    From Gopher · §Method
  • Megatron-Turing NLG2022 · cited 1×, 1 in Method
    “Pipeline model parallelism (or, pipeline parallelism) divides the layers of the model into stages that can be processed in parallel [23, 42].”
    From Megatron-Turing NLG · §Large Model Training Infrastructure
  • PaLM2022 · cited 3×
    “Therefore, techniques have arisen for splitting model tensors across accelerators (Shazeer et al. 2018) or alternatively separating layers of the models across accelerators and then pipe-lining activations between the st…”
    From PaLM · §Related Work
Abstract

Scaling up deep neural network capacity has been known as an effective approach to improving model quality for several different machine learning tasks. In many cases, increasing model capacity beyond the memory limit of a single accelerator has required developing special algorithms or infrastructure. These solutions are often architecture-specific and do not transfer to other tasks. To address the need for efficient and task-independent model parallelism, we introduce GPipe, a pipeline parallelism library that allows scaling any network that can be expressed as a sequence of layers. By pipelining different sub-sequences of layers on separate accelerators, GPipe provides the flexibility of scaling a variety of different networks to gigantic sizes efficiently. Moreover, GPipe utilizes a novel batch-splitting pipelining algorithm, resulting in almost linear speedup when a model is partitioned across multiple accelerators. We demonstrate the advantages of GPipe by training large-scale neural networks on two different tasks with distinct network architectures: (i) Image Classification: We train a 557-million-parameter AmoebaNet model and attain a top-1 accuracy of 84.4% on ImageNet-2012, (ii) Multilingual Neural Machine Translation: We train a single 6-billion-parameter, 128-layer Transformer model on a corpus spanning over 100 languages and achieve better quality than all bilingual models.