optimizationIntroduced by GShard · 2020
Automatic sharding
Annotate a few tensors and let the compiler shard the whole computation.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that GShard cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Split layers across devices and stream micro-batches through them.
Split each layer's matrices across GPUs so one model can be far larger than one device.
Papers using this
- 2021Switch Transformer
- 2021GLaM