Distributed training & systems
Parallelism and systems tricks for training models too big for one accelerator.
| Concept | Introduced by | Years | Papers |
|---|---|---|---|
| Pipeline parallelism Split layers across devices and stream micro-batches through them. | GPipe | 2018–2022 | 3 |
| Tensor (intra-layer) model parallelism Split each layer's matrices across GPUs so one model can be far larger than one device. | Megatron-LM | 2019–2022 | 6 |
| Automatic sharding Annotate a few tensors and let the compiler shard the whole computation. | GShard | 2020–2021 | 2 |
| 3D parallelism Combine data, tensor and pipeline parallelism to train 500B-parameter models. | Megatron-Turing NLG | 2022 | 0 |
| Pathways multi-pod training Train one dense model across thousands of chips in multiple TPU pods. | PaLM | 2022 | 1 |