architectureIntroduced by GShard · 2020
Sparse mixture-of-experts layers
Route each token to a few of many expert feed-forward layers: huge capacity, constant compute.
Drafted by AI · not yet reviewed
How this idea evolved
No earlier ideas recorded for this concept yet.
Papers using this
- 2015Knowledge Distillation
- 2021Switch Transformer
- 2021Gopher
- 2021GLaM
- 2021Fairseq MoE LMs
- 2022Chinchilla