Paper Lineage
Esc
MethodSep 2014arXiv 1409.0473cs.CL

Neural Machine Translation by Jointly Learning to Align and Translate

Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio

Neural machine translation is a recently proposed approach to machine translation. Unlike the traditional statistical machine translation, the neural machine translation aims at building a single neural network that can be jointly tuned to maximize the translation performance.

From the abstract

Built on

1 paper · 0 verifiedSee as graph

Also cited · not yet reviewed (1)

  • Seq2Seq2014 · cited 10×, 7 in Method
    “Sutskever et al. 2014 reported that the neural machine translation based on RNNs with long short-term memory (LSTM) units achieves close to the state-of-the-art performance of the conventional phrase-based machine transl…”
    From this paper · §Background: Neural Machine Translation

Led to

  • Seq2Seq2014 · cited 4×
    “We were initially convinced that the LSTM would fail on long sentences due to its limited memory, and other researchers reported poor performance on long sentences with a model similar to ours [5, 2, 26].”
    From Seq2Seq · §Conclusion
  • GNMT2016 · cited 6×, 3 in Method
    “Our attention module is similar to [2].”
    From GNMT · §Model Architecture
  • ByteNet2016 · cited 4×, 1 in Method
    “In our comparison we consider the following neural translation models: the Recurrent Continuous Translation Model (RCTM) 1 and 2 (Kalchbrenner & Blunsom 2013); the RNN Enc-Dec (Sutskever et al. 2014; Cho et al. 2014); th…”
    From ByteNet · §Model Comparison
  • ConvS2S2017 · cited 6×
    “The dominant approach to date encodes the input sequence with a series of bi-directional recurrent neural networks (RNN) and generates a variable length output with another set of decoder RNNs, both of which interface vi…”
    From ConvS2S · §Introduction
  • Transformer2017 · cited 5×, 3 in Method
    “Most competitive neural sequence transduction models have an encoder-decoder structure [5, 2, 35].”
    From Transformer · §Model Architecture
  • CoQA2018 · cited 1×, 1 in Method
    “Motivated by their success, we use a sequence-to-sequence with attention model for generating answers Bahdanau et al. 2015.”
    From CoQA · §Models
  • LXMERT2019 · cited 1×, 1 in Method
    “Attention layers Bahdanau et al. 2014; Xu et al. 2015 aim to retrieve information from a set of context vectors {yj}\{y_{j}\} related to a query vector xx.”
    From LXMERT · §Model Architecture
  • T52019 · cited 2×, 1 in Method
    “Self-attention is a variant of attention (Graves 2013; Bahdanau et al. 2015) that processes a sequence by replacing each element by a weighted average of the rest of the sequence.”
    From T5 · §Setup
  • Perceiver2021 · cited 1×, 1 in Method
    “Both cross-attention and Transformer modules are structured around the use of query-key-value (QKV) attention (Graves et al. 2014; Weston et al. 2015; Bahdanau et al. 2015).”
    From Perceiver · §Methods
Abstract

Neural machine translation is a recently proposed approach to machine translation. Unlike the traditional statistical machine translation, the neural machine translation aims at building a single neural network that can be jointly tuned to maximize the translation performance. The models proposed recently for neural machine translation often belong to a family of encoder-decoders and consists of an encoder that encodes a source sentence into a fixed-length vector from which a decoder generates a translation. In this paper, we conjecture that the use of a fixed-length vector is a bottleneck in improving the performance of this basic encoder-decoder architecture, and propose to extend this by allowing a model to automatically (soft-)search for parts of a source sentence that are relevant to predicting a target word, without having to form these parts as a hard segment explicitly. With this new approach, we achieve a translation performance comparable to the existing state-of-the-art phrase-based system on the task of English-to-French translation. Furthermore, qualitative analysis reveals that the (soft-)alignments found by the model agree well with our intuition.