GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Neural network scaling has been critical for improving the model quality in many real-world machine learning applications with vast amounts of training data and compute. Although this trend of scaling is affirmed to be a sure-fire approach for better model quality, there are challenges on the path such as the computation cost, ease of programming, and efficient implementation on parallel devices.
Also cited · not yet reviewed (6)
- GPipe2018 · cited 15דSimilarly in natural language processing, scaling Transformers [10] yielded consistent gains on language understanding tasks [4, 11, 12], cross-lingual down-stream transfer [4, 13] and (massively-)multilingual neural mac…”From this paper · §Introduction
- BERT2018 · cited 5דFor years, the fields have been continuously reporting new state of the art results using varieties of model architectures for computer vision tasks [57, 58, 7], for natural language understanding tasks [59, 60, 61], for…”From this paper · §Related Work
- Kaplan scaling laws2020 · cited 4דScaling neural networks brings dramatic quality gains over a wide array of machine learning problems [1, 2, 3, 4, 5, 6].”From this paper · §Introduction
- GPT-32020 · cited 3דScaling neural networks brings dramatic quality gains over a wide array of machine learning problems [1, 2, 3, 4, 5, 6].”From this paper · §Introduction
Show 2 more
- ResNet2015 · cited 2דFor years, the fields have been continuously reporting new state of the art results using varieties of model architectures for computer vision tasks [57, 58, 7], for natural language understanding tasks [59, 60, 61], for…”From this paper · §Related Work
- T52019 · cited 2דWithout loss of generality, emerging approaches follow scaling Transformer by stacking more and more layers [49, 15], widening the governing dimensions of the network (i.e. model dimension, hidden dimension or number of…”From this paper · §Massively Multilingual, Massive Machine Translation (M4)
Led to
- Switch Transformer2021 · cited 5דMoE models have had notable successes in machine translation (Shazeer et al. 2017; Shazeer et al. 2018; Lepikhin et al. 2020), however, widespread adoption is hindered by complexity, communication costs, and training ins…”From Switch Transformer · §Introduction
- GLaM2021 · cited 4×, 2 in Method“Similar to the GShard MoE Transformer (Lepikhin et al. 2021), we replace the feed-forward component of every other Transformer layer with an MoE layer, as shown in Figure 2.”From GLaM · §Model Architecture
- Fairseq MoE LMs2021 · cited 7×, 4 in Method“Our MoE models follow the design proposed in Lepikhin et al. 2021 with alternating dense and expert layers and top-2 expert selection.”From Fairseq MoE LMs · §Experimental Setup
- PaLM2022 · cited 2דMany other works aim to increase of the scale of models, while limiting communication overheads (Rajbhandari et al. 2020; Lepikhin et al. 2020; Li et al. 2020; Rasley et al. 2020; Rajbhandari et al. 2021; Ren et al. 2021…”From PaLM · §Related Work
Abstract
Neural network scaling has been critical for improving the model quality in many real-world machine learning applications with vast amounts of training data and compute. Although this trend of scaling is affirmed to be a sure-fire approach for better model quality, there are challenges on the path such as the computation cost, ease of programming, and efficient implementation on parallel devices. GShard is a module composed of a set of lightweight annotation APIs and an extension to the XLA compiler. It provides an elegant way to express a wide range of parallel computation patterns with minimal changes to the existing model code. GShard enabled us to scale up multilingual neural machine translation Transformer model with Sparsely-Gated Mixture-of-Experts beyond 600 billion parameters using automatic sharding. We demonstrate that such a giant model can efficiently be trained on 2048 TPU v3 accelerators in 4 days to achieve far superior quality for translation from 100 languages to English compared to the prior art.