GLU Variants Improve Transformer
Noam Shazeer
08083) consist of the component-wise product of two linear projections, one of which is first passed through a sigmoid function. Variations on GLU are possible, using different nonlinear (or even linear) functions in place of sigmoid.
From the abstract
Also cited · not yet reviewed (2)
- T52019 · cited 10דFollowing the T5 codebase [Raffel et al. 2019] 11 1 Also in the interest of ML fairness., we use a version with no bias:”From this paper · §Introduction
- Transformer2017 · cited 2דThe Transformer [Vaswani et al. 2017] sequence-to-sequence model alternates between multi-head attention, and what it calls "position-wise feed-forward networks" (FFN).”From this paper · §Introduction
Led to
- LaMDA2022 · cited 1×, 1 in Method“The Transformer has 64 layers, dmodel=8192d_{model}=8192, dff=65536d_{ff}=65536, h=128h=128, dk=dv=128d_{k}=d_{v}=128, relative attention as described in T5 [11], and gated-GELU activation as described in Raffel et…”From LaMDA · §LaMDA pre-training
- PaLM2022 · cited 2×, 2 in Method“SwiGLU Activation – We use SwiGLU activations (Swish(xW)⋅xV\textrm{Swish}(xW)\cdot xV) for the MLP intermediate activations because they have been shown to significantly increase quality compared to standard ReLU, GeL…”From PaLM · §Model Architecture
Abstract
Gated Linear Units (arXiv:1612.08083) consist of the component-wise product of two linear projections, one of which is first passed through a sigmoid function. Variations on GLU are possible, using different nonlinear (or even linear) functions in place of sigmoid. We test these variants in the feed-forward sublayers of the Transformer (arXiv:1706.03762) sequence-to-sequence model, and find that some of them yield quality improvements over the typically-used ReLU or GELU activations.