Paper Lineage
Esc
MethodMay 2017arXiv 1705.03122cs.CL

Convolutional Sequence to Sequence Learning

Jonas Gehring, Michael Auli, David Grangier and 2 others

The prevalent approach to sequence to sequence learning maps an input sequence to a variable length output sequence via recurrent neural networks. We introduce an architecture based entirely on convolutional neural networks.

From the abstract

Built on

3 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (3)

  • Bahdanau attention2014 · cited 6×
    “The dominant approach to date encodes the input sequence with a series of bi-directional recurrent neural networks (RNN) and generates a variable length output with another set of decoder RNNs, both of which interface vi…”
    From this paper · §Introduction
  • ByteNet2016 · cited 5×
    “On WMT’14 English to German translation we compare to the following prior work: Luong et al. (2015) is based on a four layer LSTM attention model, ByteNet (Kalchbrenner et al., 2016) propose a convolutional model based o…”
    From this paper · §Results
  • Seq2Seq2014 · cited 3×
    “Sequence to sequence learning has been successful in many tasks such as machine translation, speech recognition (Sutskever et al., 2014; Chorowski et al., 2015) and text summarization (Rush et al., 2015; Nallapati et al.…”
    From this paper · §Introduction

Led to

  • Transformer2017 · cited 6×, 5 in Method
    “In the following sections, we will describe the Transformer, motivate self-attention and discuss its advantages over models such as [17, 18] and [9].”
    From Transformer · §Background
  • Perceiver2021 · cited 1×, 1 in Method
    “The latent array itself is initialized using a learned position encoding (Gehring et al. 2017) (see Appendix Sec.”
    From Perceiver · §Methods
  • RoFormer (RoPE)2021 · cited 2×
    “Convolution neural networks (CNNs) based models (CNNs) Gehring et al. 2017 were typically considered position-agnostic, but recent work Islam et al. 2020 has shown that the commonly used padding operation can implicitly…”
    From RoFormer (RoPE) · §Introduction
Abstract

The prevalent approach to sequence to sequence learning maps an input sequence to a variable length output sequence via recurrent neural networks. We introduce an architecture based entirely on convolutional neural networks. Compared to recurrent models, computations over all elements can be fully parallelized during training and optimization is easier since the number of non-linearities is fixed and independent of the input length. Our use of gated linear units eases gradient propagation and we equip each decoder layer with a separate attention module. We outperform the accuracy of the deep LSTM setup of Wu et al. (2016) on both WMT'14 English-German and WMT'14 English-French translation at an order of magnitude faster speed, both on GPU and CPU.