Paper Lineage
Esc
MethodJun 2017arXiv 1706.03762cs.CL

Attention Is All You Need

Ashish Vaswani, Noam Shazeer, Niki Parmar and 5 others

The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism.

From the abstract

Built on

6 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (6)

  • ConvS2S2017 · cited 6×, 5 in Method
    “In the following sections, we will describe the Transformer, motivate self-attention and discuss its advantages over models such as [17, 18] and [9].”
    From this paper · §Background
  • Bahdanau attention2014 · cited 5×, 3 in Method
    “Most competitive neural sequence transduction models have an encoder-decoder structure [5, 2, 35].”
    From this paper · §Model Architecture
  • ByteNet2016 · cited 3×, 2 in Method
    “In the following sections, we will describe the Transformer, motivate self-attention and discuss its advantages over models such as [17, 18] and [9].”
    From this paper · §Background
  • GNMT2016 · cited 6×, 1 in Method
    “This mimics the typical encoder-decoder attention mechanisms in sequence-to-sequence models such as [38, 2, 9].”
    From this paper · §Model Architecture
Show 2 more
  • Seq2Seq2014 · cited 2×, 1 in Method
    “Most competitive neural sequence transduction models have an encoder-decoder structure [5, 2, 35].”
    From this paper · §Model Architecture
  • ResNet2015 · cited 1×, 1 in Method
    “We employ a residual connection [11] around each of the two sub-layers, followed by layer normalization [1].”
    From this paper · §Model Architecture

Led to

  • decaNLP2018 · cited 4×
    “We provide a set of baselines for decaNLP that combine the basics of sequence-to-sequence learning [Sutskever et al. 2014, Bahdanau et al. 2014, Luong et al. 2015b] with pointer networks [Vinyals et al. 2015, Merity et a…”
    From decaNLP · §Introduction
  • Set Transformer2018 · cited 3×, 1 in Method
    “Transformer, (Vaswani et al. 2017)).”
    From Set Transformer · §Introduction
  • BERT2018 · cited 4×
    “BERT’s model architecture is a multi-layer bidirectional Transformer encoder based on the original implementation described in Vaswani et al. 2017 and released in the tensor2tensor library.11 1 https://github.com/tensorf…”
    From BERT · §BERT
  • GPipe2018 · cited 4×
    “Our comparison is based on the performance of a single Transformer [15] trained on all language pairs in this corpus.”
    From GPipe · §Massive Massively Multilingual Machine Translation
  • Transformer-XL2019 · cited 5×, 3 in Method
    “Firstly, in the standard Transformer (Vaswani et al. 2017), the attention score between query qiq_{i} and key vector kjk_{j} within the same segment can be decomposed as”
    From Transformer-XL · §Model
  • Adapters2019 · cited 3×
    “Recent state-of-the-art results on question answering (Rajpurkar et al. 2016) and text classification (Wang et al. 2018) have been attained by fine-tuning a Transformer network (Vaswani et al. 2017) with a Masked Languag…”
    From Adapters · §Related Work
  • UniLM2019 · cited 3×, 1 in Method
    “As shown in Figure 1, the pre-training optimizes the shared Transformer [43] network with respect to several unsupervised language modeling objectives, namely, unidirectional LM, bidirectional LM, and sequence-to-sequenc…”
    From UniLM · §Unified Language Model Pre-training
  • XLNet2019 · cited 1×, 1 in Method
    “where Q, K, V denote the query, key, and value in an attention operation [33].”
    From XLNet · §Proposed Method
  • RoBERTa2019 · cited 1×, 1 in Method
    “BERT uses the now ubiquitous transformer architecture Vaswani et al. 2017, which we will not review in detail.”
    From RoBERTa · §Background
  • ViLBERT2019 · cited 2×, 2 in Method
    “These tokens are mapped to learned encodings and passed through LL “encoder-style” transformer blocks [27] to produce final representations h0,…,hTh_{0},\dots,h_{T}.”
    From ViLBERT · §Approach
  • VisualBERT2019 · cited 2×, 1 in Method
    “BERT (Devlin et al. 2019) is a Transformer (Vaswani et al. 2017) with subwords (Wu et al. 2016) as input and trained using language modeling objectives.”
    From VisualBERT · §A Joint Representation Model for Vision and Language
  • LXMERT2019 · cited 6×, 4 in Method
    “We build our cross-modality model with self-attention and cross-attention layers following the recent progress in designing natural language processing models (e.g., transformers Vaswani et al. 2017).”
    From LXMERT · §Model Architecture
  • VL-BERT2019 · cited 5×
    “In the milestone work of Transformers (Vaswani et al. 2017), the Transformer attention module is proposed as a generic building block for various NLP tasks.”
    From VL-BERT · §Related Work
  • Megatron-LM2019 · cited 3×, 3 in Method
    “Current work in NLP trends towards using transformer models (Vaswani et al. 2017) due to their superior accuracy and compute efficiency.”
    From Megatron-LM · §Background and Challenges
  • RLHF for LMs (Ziegler)2019 · cited 1×, 1 in Method
    “The model is a Transformer with 36 layers, 20 heads, and embedding size 1280 [Vaswani et al. 2017].”
    From RLHF for LMs (Ziegler) · §Methods
  • ALBERT2019 · cited 2×
    “The backbone of the ALBERT architecture is similar to BERT in that it uses a transformer encoder (Vaswani et al. 2017) with GELU nonlinearities (Hendrycks & Gimpel 2016).”
    From ALBERT · §The Elements of ALBERT
  • T52019 · cited 6×, 5 in Method
    “Early results on transfer learning for NLP leveraged recurrent neural networks (Peters et al. 2018; Howard and Ruder 2018), but it has recently become more common to use models based on the “Transformer” architecture (Va…”
    From T5 · §Setup
  • BART2019 · cited 2×, 2 in Method
    “BART uses the standard sequence-to-sequence Transformer architecture from Vaswani et al. 2017, except, following GPT, that we modify ReLU activation functions to GeLUs (Hendrycks & Gimpel 2016) and initialise parameters…”
    From BART · §Model
  • 12-in-12019 · cited 1×, 1 in Method
    “Each stream is a series of transformer blocks (TRM) vaswani2017attention connected by co-attentional transformer layers (Co-TRM) which enable information exchange between modalities.”
    From 12-in-1 · §Approach
  • Meshed-Memory Transformer2019 · cited 7×
    “The recent advent of fully-attentive models, in which the recurrent relation is abandoned in favour of the use of self-attention, offers unique opportunities in terms of set and sequence modeling performances, as testifi…”
    From Meshed-Memory Transformer · §Introduction
  • Kaplan scaling laws2020 · cited 2×, 1 in Method
    “In this work we will empirically investigate the dependence of language modeling loss on all of these factors, focusing on the Transformer architecture [VSP+17, LSP+18].”
    From Kaplan scaling laws · §Introduction
  • Knowledge in LM parameters2020 · cited 1×, 1 in Method
    “Currently, the most popular model architectures used in transfer learning for NLP are Transformer-based Vaswani et al. 2017 “encoder-only” models like BERT Devlin et al. 2018.”
    From Knowledge in LM parameters · §Background
  • GLU variants2020 · cited 2×
    “The Transformer [Vaswani et al. 2017] sequence-to-sequence model alternates between multi-head attention, and what it calls "position-wise feed-forward networks" (FFN).”
    From GLU variants · §Introduction
  • X-Linear attention2020 · cited 2×
    “Note that here each key/value is concatenated with the new attended feature, followed with a residual connection and layer normalization as in vaswani2017attention.”
    From X-Linear attention · §X-linear Attention Networks (X-LAN)
  • GPT-32020 · cited 2×
    “Work in this vein has successively increased model size: 213 million parameters [134] in the original paper, 300 million parameters [20], 1.5 billion parameters [117], 8 billion parameters [125], 11 billion parameters [1…”
    From GPT-3 · §Related Work
  • “Our work is most similar to [73], who also train Transformer models [62] to optimize human feedback across a range of tasks, including summarization on the Reddit TL;DR and CNN/DM datasets.”
    From Learning to summarize from human feedbac · §unknown section
  • ViT2020 · cited 4×, 2 in Method
    “In model design we follow the original Transformer (Vaswani et al. 2017) as closely as possible.”
    From ViT · §Method
  • VL-BERT meta-analysis2020 · cited 2×, 2 in Method
    “This is similar to the input masking applied in autoregressive Transformer decoders (Vaswani et al. 2017).”
    From VL-BERT meta-analysis · §A Unified Framework
  • DeiT2020 · cited 7×
    “introduced by Vaswani et al. [52] for machine translation are currently the reference model for all natural language processing (NLP) tasks.”
    From DeiT · §Related work
  • Switch Transformer2021 · cited 2×
    “The guiding design principle for Switch Transformers is to maximize the parameter count of a Transformer model (Vaswani et al. 2017) in a simple and computationally efficient way.”
    From Switch Transformer · §Switch Transformer
  • VL-T52021 · cited 3×, 2 in Method
    “We use transformer encoder-decoder architecture (Vaswani et al. 2017) to encode visual and text inputs and generate label text.”
    From VL-T5 · §Model
  • Conceptual 12M2021 · cited 2×, 1 in Method
    “For ic-based pre-training and downstream tasks, we follow the state-of-the-art architecture that heavily rely on self-attention [78] or similar mechanisms [70, 85, 19, 33, 23].”
    From Conceptual 12M · §Evaluating Vision-and-Language Pre-Training Data
  • DALL·E2021 · cited 2×, 1 in Method
    “Our goal is to train a transformer (Vaswani et al. 2017) to autoregressively model the text and image tokens as a single stream of data.”
    From DALL·E · §Method
  • CLIP2021 · cited 1×, 1 in Method
    “The text encoder is a Transformer (Vaswani et al. 2017) with the architecture modifications described in Radford et al. 2019.”
    From CLIP · §Approach
  • Perceiver2021 · cited 7×, 3 in Method
    “Our latent Transformer uses the GPT-2 architecture (Radford et al. 2019), which itself is based on the decoder of the original Transformer architecture (Vaswani et al. 2017).”
    From Perceiver · §Methods
  • Swin2021 · cited 3×, 1 in Method
    “The standard Transformer architecture [64] and its adaptation for image classification [20] both conduct global self-attention, where the relationships between a token and all other tokens are computed.”
    From Swin · §Method
  • RoFormer (RoPE)2021 · cited 8×, 3 in Method
    “Generally speaking, all these approaches attempt to modify Equation 6 based on the decomposition of Equation 3 under the self-attention settings in Equation 2, which was originally proposed in Vaswani et al. 2017.”
    From RoFormer (RoPE) · §Background and Related Work
  • BEiT2021 · cited 2×, 1 in Method
    “Following ViT [12], we use the standard Transformer [42] as the backbone network.”
    From BEiT · §Methods
  • Dynamic Head2021 · cited 1×, 1 in Method
    “Recently, there is a trend to introduce the Transformer module [29] from natural language processing into computer vision tasks.”
    From Dynamic Head · §Our Approach
  • Frozen2021 · cited 2×, 1 in Method
    “Our method starts from a pre-trained deep auto-regressive language model, based on the Transformer architecture [40, 29], which parametrizes a probability distribution over text 𝐲y.”
    From Frozen · §The Frozen Method
  • ALBEF2021 · cited 1×, 1 in Method
    “We use a 6-layer transformer [39] for both the text encoder and the multimodal encoder.”
    From ALBEF · §ALBEF Pre-training
  • ViT-VQGAN2021 · cited 2×
    “Driven by the effectiveness of VQVAE and progress in sequence modeling (Vaswani et al. 2017; Devlin et al. 2019), many approaches follow the two-stage paradigm.”
    From ViT-VQGAN · §Related Work
  • VLMo2021 · cited 2×
    “Pre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”
    From VLMo · §Related Work
  • MAE2021 · cited 3×
    “Motivated by the success in NLP, related recent methods Chen2020c; Dosovitskiy2021; Bao2021 are based on Transformers Vaswani2017. iGPT Chen2020c operates on sequences of pixels and predicts unknown pixels.”
    From MAE · §Related Work
  • Swin V22021 · cited 3×
    “Transformer has served the standard network since the pioneer work of vaswani2017attention.”
    From Swin V2 · §Related Works
  • Florence2021 · cited 1×, 1 in Method
    “Our Florence pretrained model uses a two-tower architecture: a 12-layer transformer (Vaswani et al. 2017) as language encoder, similar to CLIP (Radford et al. 2021), and a hierarchical Vision Transformer as the image enc…”
    From Florence · §Approach
  • Gopher2021 · cited 2×, 1 in Method
    “Progress has been driven by both scale and network architecture (Hochreiter and Schmidhuber 1997; Bahdanau et al. 2014; Vaswani et al. 2017). rosenfeld2020a and Kaplan et al. 2020 independently found power laws relating…”
    From Gopher · §Introduction
  • Fairseq MoE LMs2021 · cited 1×, 1 in Method
    “While numerous variations have been proposed, such LMs are predominantly based on the transformer architecture (Vaswani et al. 2017).”
    From Fairseq MoE LMs · §Background and Related Work
  • LaMDA2022 · cited 1×, 1 in Method
    “We use a decoder-only Transformer [92] language model as the model architecture for LaMDA.”
    From LaMDA · §LaMDA pre-training
  • PaLM2022 · cited 3×, 1 in Method
    “PaLM uses a standard Transformer model architecture (Vaswani et al. 2017) in a decoder-only setup (i.e., each timestep can only attend to itself and past timesteps), with the following modifications:”
    From PaLM · §Model Architecture
  • Flamingo2022 · cited 2×
    “Language modelling has recently made substantial progress following the introduction of Transformers [115].”
    From Flamingo · §Related work
  • BIG-bench2022 · cited 2×
    “We use 13 dense decoder-only Transformer models (Vaswani et al. 2017) with gated activation layers (Dauphin et al. 2017) and GELU activations based on the LaMDA architectures (Thoppilan et al. 2022).”
    From BIG-bench · §What is in BIG-bench?
Abstract

The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles by over 2 BLEU. On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.8 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature. We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data.