Paper Lineage
Esc

Distributed training & systems

Parallelism and systems tricks for training models too big for one accelerator.

Concepts in Distributed training & systems
ConceptIntroduced byYearsPapers
Pipeline parallelism
Split layers across devices and stream micro-batches through them.
GPipe2018–20223
Tensor (intra-layer) model parallelism
Split each layer's matrices across GPUs so one model can be far larger than one device.
Megatron-LM2019–20226
Automatic sharding
Annotate a few tensors and let the compiler shard the whole computation.
GShard2020–20212
3D parallelism
Combine data, tensor and pipeline parallelism to train 500B-parameter models.
Megatron-Turing NLG20220
Pathways multi-pod training
Train one dense model across thousands of chips in multiple TPU pods.
PaLM20221