Paper Lineage
Esc
MethodJun 2020arXiv 2006.16668cs.CL

GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu and 6 others

Neural network scaling has been critical for improving the model quality in many real-world machine learning applications with vast amounts of training data and compute. Although this trend of scaling is affirmed to be a sure-fire approach for better model quality, there are challenges on the path such as the computation cost, ease of programming, and efficient implementation on parallel devices.

From the abstract

Built on

6 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (6)

  • GPipe2018 · cited 15×
    “Similarly in natural language processing, scaling Transformers [10] yielded consistent gains on language understanding tasks [4, 11, 12], cross-lingual down-stream transfer [4, 13] and (massively-)multilingual neural mac…”
    From this paper · §Introduction
  • BERT2018 · cited 5×
    “For years, the fields have been continuously reporting new state of the art results using varieties of model architectures for computer vision tasks [57, 58, 7], for natural language understanding tasks [59, 60, 61], for…”
    From this paper · §Related Work
  • Kaplan scaling laws2020 · cited 4×
    “Scaling neural networks brings dramatic quality gains over a wide array of machine learning problems [1, 2, 3, 4, 5, 6].”
    From this paper · §Introduction
  • GPT-32020 · cited 3×
    “Scaling neural networks brings dramatic quality gains over a wide array of machine learning problems [1, 2, 3, 4, 5, 6].”
    From this paper · §Introduction
Show 2 more
  • ResNet2015 · cited 2×
    “For years, the fields have been continuously reporting new state of the art results using varieties of model architectures for computer vision tasks [57, 58, 7], for natural language understanding tasks [59, 60, 61], for…”
    From this paper · §Related Work
  • T52019 · cited 2×
    “Without loss of generality, emerging approaches follow scaling Transformer by stacking more and more layers [49, 15], widening the governing dimensions of the network (i.e. model dimension, hidden dimension or number of…”
    From this paper · §Massively Multilingual, Massive Machine Translation (M4)

Led to

  • Switch Transformer2021 · cited 5×
    “MoE models have had notable successes in machine translation (Shazeer et al. 2017; Shazeer et al. 2018; Lepikhin et al. 2020), however, widespread adoption is hindered by complexity, communication costs, and training ins…”
    From Switch Transformer · §Introduction
  • GLaM2021 · cited 4×, 2 in Method
    “Similar to the GShard MoE Transformer (Lepikhin et al. 2021), we replace the feed-forward component of every other Transformer layer with an MoE layer, as shown in Figure 2.”
    From GLaM · §Model Architecture
  • Fairseq MoE LMs2021 · cited 7×, 4 in Method
    “Our MoE models follow the design proposed in Lepikhin et al. 2021 with alternating dense and expert layers and top-2 expert selection.”
    From Fairseq MoE LMs · §Experimental Setup
  • PaLM2022 · cited 2×
    “Many other works aim to increase of the scale of models, while limiting communication overheads (Rajbhandari et al. 2020; Lepikhin et al. 2020; Li et al. 2020; Rasley et al. 2020; Rajbhandari et al. 2021; Ren et al. 2021…”
    From PaLM · §Related Work
Abstract

Neural network scaling has been critical for improving the model quality in many real-world machine learning applications with vast amounts of training data and compute. Although this trend of scaling is affirmed to be a sure-fire approach for better model quality, there are challenges on the path such as the computation cost, ease of programming, and efficient implementation on parallel devices. GShard is a module composed of a set of lightweight annotation APIs and an extension to the XLA compiler. It provides an elegant way to express a wide range of parallel computation patterns with minimal changes to the existing model code. GShard enabled us to scale up multilingual neural machine translation Transformer model with Sparsely-Gated Mixture-of-Experts beyond 600 billion parameters using automatic sharding. We demonstrate that such a giant model can efficiently be trained on 2048 TPU v3 accelerators in 4 days to achieve far superior quality for translation from 100 languages to English compared to the prior art.