Paper Lineage
Esc
optimizationIntroduced by Megatron-LM · 2019

Tensor (intra-layer) model parallelism

Split each layer's matrices across GPUs so one model can be far larger than one device.

Drafted by AI · not yet reviewed

How this idea evolved

Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that Megatron-LM cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.

Papers using this