optimizationIntroduced by Megatron-Turing NLG · 2022
3D parallelism
Combine data, tensor and pipeline parallelism to train 500B-parameter models.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that Megatron-Turing NLG cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Split each layer's matrices across GPUs so one model can be far larger than one device.
Split layers across devices and stream micro-batches through them.
Papers using this
No other paper in this dataset is tagged with it yet.