optimizationIntroduced by PaLM · 2022
Pathways multi-pod training
Train one dense model across thousands of chips in multiple TPU pods.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that PaLM cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Combine data, tensor and pipeline parallelism to train 500B-parameter models.
Split layers across devices and stream micro-batches through them.
Annotate a few tensors and let the compiler shard the whole computation.
Papers using this
- 2022BIG-bench