Paper Lineage
Esc
MethodFeb 2020arXiv 2002.05202cs.LG

GLU Variants Improve Transformer

Noam Shazeer

08083) consist of the component-wise product of two linear projections, one of which is first passed through a sigmoid function. Variations on GLU are possible, using different nonlinear (or even linear) functions in place of sigmoid.

From the abstract

Built on

2 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (2)

  • T52019 · cited 10×
    “Following the T5 codebase [Raffel et al. 2019] 11 1 Also in the interest of ML fairness., we use a version with no bias:”
    From this paper · §Introduction
  • Transformer2017 · cited 2×
    “The Transformer [Vaswani et al. 2017] sequence-to-sequence model alternates between multi-head attention, and what it calls "position-wise feed-forward networks" (FFN).”
    From this paper · §Introduction

Led to

  • LaMDA2022 · cited 1×, 1 in Method
    “The Transformer has 64 layers, dm​o​d​e​l=8192d_{model}=8192, df​f=65536d_{ff}=65536, h=128h=128, dk=dv=128d_{k}=d_{v}=128, relative attention as described in T5 [11], and gated-GELU activation as described in Raffel et…”
    From LaMDA · §LaMDA pre-training
  • PaLM2022 · cited 2×, 2 in Method
    “SwiGLU Activation – We use SwiGLU activations (Swish​(x​W)⋅x​V\textrm{Swish}(xW)\cdot xV) for the MLP intermediate activations because they have been shown to significantly increase quality compared to standard ReLU, GeL…”
    From PaLM · §Model Architecture
Abstract

Gated Linear Units (arXiv:1612.08083) consist of the component-wise product of two linear projections, one of which is first passed through a sigmoid function. Variations on GLU are possible, using different nonlinear (or even linear) functions in place of sigmoid. We test these variants in the feed-forward sublayers of the Transformer (arXiv:1706.03762) sequence-to-sequence model, and find that some of them yield quality improvements over the typically-used ReLU or GELU activations.