Sparse & mixture-of-experts models
Conditional computation: activate only part of a huge model for each input.
| Concept | Introduced by | Years | Papers |
|---|---|---|---|
| Sparse mixture-of-experts layers Route each token to a few of many expert feed-forward layers: huge capacity, constant compute. | GShard | 2020–2022 | 6 |
| Top-1 (switch) routing Send each token to a single expert, simplifying MoE and scaling to trillions of parameters. | Switch Transformer | 2021–2022 | 1 |