Paper Lineage
Esc
MethodSep 2016arXiv 1609.08144cs.CL

Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

Yonghui Wu, Mike Schuster, Zhifeng Chen and 28 others

Neural Machine Translation (NMT) is an end-to-end learning approach for automated translation, with the potential to overcome many of the weaknesses of conventional phrase-based translation systems. Unfortunately, NMT systems are known to be computationally expensive both in training and in translation inference.

From the abstract

Built on

3 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (3)

  • Bahdanau attention2014 · cited 6×, 3 in Method
    “Our attention module is similar to [2].”
    From this paper · §Model Architecture
  • Seq2Seq2014 · cited 6×, 2 in Method
    “This observation is similar to previous observations that deep LSTMs significantly outperform shallow LSTMs [41].”
    From this paper · §Model Architecture
  • ResNet2015 · cited 2×, 1 in Method
    “Motivated by the idea of modeling differences between an intermediate layer’s output and the targets, which has shown to work well for many projects in the past [16, 21, 40], we introduce residual connections among the L…”
    From this paper · §Model Architecture

Led to

  • ByteNet2016 · cited 2×
    “On the character-level machine translation task, ByteNet betters a comparable version of GNMT (Wu et al. 2016a) that is a state-of-the-art system.”
    From ByteNet · §Introduction
  • Transformer2017 · cited 6×, 1 in Method
    “This mimics the typical encoder-decoder attention mechanisms in sequence-to-sequence models such as [38, 2, 9].”
    From Transformer · §Model Architecture
  • SentencePiece2018 · cited 3×
    “Neural machine translation (NMT) Bahdanau et al. 2014; Luong et al. 2015; Wu et al. 2016; Vaswani et al. 2017 has especially gained increasing popularity, as it can leverage neural networks to directly perform translatio…”
    From SentencePiece · §Introduction
  • UniLM2019 · cited 1×, 1 in Method
    “Texts are tokenized to subword units by WordPiece [48].”
    From UniLM · §Unified Language Model Pre-training
  • ViLBERT2019 · cited 1×, 1 in Method
    “For a given token, the input representation is a sum of a token-specific learned embedding [28] and encodings for position (i.e. token’s index in the sequence) and segment (i.e. index of the token’s sentence if multiple…”
    From ViLBERT · §Approach
  • VisualBERT2019 · cited 1×, 1 in Method
    “BERT (Devlin et al. 2019) is a Transformer (Vaswani et al. 2017) with subwords (Wu et al. 2016) as input and trained using language modeling objectives.”
    From VisualBERT · §A Joint Representation Model for Vision and Language
  • Unicoder-VL2019 · cited 1×, 1 in Method
    “TT is the length of the WordPiece [\citeauthoryearWu et al.2016] linguistic input.”
    From Unicoder-VL · §Approach
  • LXMERT2019 · cited 2×, 2 in Method
    “A sentence is first split into words {w1,…,wn}\left\{w_{1},\ldots,w_{n}\right\} with length of nn by the same WordPiece tokenizer Wu et al. 2016 in Devlin et al. 2019.”
    From LXMERT · §Model Architecture
  • VL-BERT2019 · cited 2×
    “Token Embedding Following the practice in BERT, the linguistic words are embedded with WordPiece embeddings (Wu et al. 2016) with a 30,000 vocabulary.”
    From VL-BERT · §VL-BERT
  • VLMo2021 · cited 1×, 1 in Method
    “Following BERT [10], we tokenize the text to subword units by WordPiece [47].”
    From VLMo · §Methods
Abstract

Neural Machine Translation (NMT) is an end-to-end learning approach for automated translation, with the potential to overcome many of the weaknesses of conventional phrase-based translation systems. Unfortunately, NMT systems are known to be computationally expensive both in training and in translation inference. Also, most NMT systems have difficulty with rare words. These issues have hindered NMT's use in practical deployments and services, where both accuracy and speed are essential. In this work, we present GNMT, Google's Neural Machine Translation system, which attempts to address many of these issues. Our model consists of a deep LSTM network with 8 encoder and 8 decoder layers using attention and residual connections. To improve parallelism and therefore decrease training time, our attention mechanism connects the bottom layer of the decoder to the top layer of the encoder. To accelerate the final translation speed, we employ low-precision arithmetic during inference computations. To improve handling of rare words, we divide words into a limited set of common sub-word units ("wordpieces") for both input and output. This method provides a good balance between the flexibility of "character"-delimited models and the efficiency of "word"-delimited models, naturally handles translation of rare words, and ultimately improves the overall accuracy of the system. Our beam search technique employs a length-normalization procedure and uses a coverage penalty, which encourages generation of an output sentence that is most likely to cover all the words in the source sentence. On the WMT'14 English-to-French and English-to-German benchmarks, GNMT achieves competitive results to state-of-the-art. Using a human side-by-side evaluation on a set of isolated simple sentences, it reduces translation errors by an average of 60% compared to Google's phrase-based production system.