Paper Lineage
Esc
MethodJan 2019arXiv 1901.02860cs.LG

Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context

Zihang Dai, Zhilin Yang, Yiming Yang and 3 others

Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling. We propose a novel neural architecture Transformer-XL that enables learning dependency beyond a fixed length without disrupting temporal coherence.

From the abstract

Built on

3 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (3)

  • Transformer2017 · cited 5×, 3 in Method
    “Firstly, in the standard Transformer (Vaswani et al. 2017), the attention score between query qiq_{i} and key vector kjk_{j} within the same segment can be decomposed as”
    From this paper · §Model
  • ELMo2018 · cited 2×, 1 in Method
    “Second, though it is possible to use padding to respect the sentence or other semantic boundaries, in practice it has been standard practice to simply chunk long text into fixed-length segments due to improved efficiency…”
    From this paper · §Model
  • BERT2018 · cited 2×, 1 in Method
    “Second, though it is possible to use padding to respect the sentence or other semantic boundaries, in practice it has been standard practice to simply chunk long text into fixed-length segments due to improved efficiency…”
    From this paper · §Model

Led to

  • XLNet2019 · cited 4×, 2 in Method
    “Inspired by the latest advancements in AR language modeling, XLNet integrates the segment recurrence mechanism and relative encoding scheme of Transformer-XL [9] into pretraining, which empirically improves the performan…”
    From XLNet · §Introduction
  • Megatron-LM2019 · cited 2×, 1 in Method
    “Recent parallel work (Ramachandran et al. 2016; Howard & Ruder 2018; Radford et al. 2018; Devlin et al. 2018; Liu et al. 2019b; Dai et al. 2019; Yang et al. 2019; Liu et al. 2019a; Lan et al. 2019) further builds upon th…”
    From Megatron-LM · §Background and Challenges
  • RoFormer (RoPE)2021 · cited 2×, 1 in Method
    “Keeping the form of Equation 3, the authors Dai et al. 2019 have proposed to decompose 𝒒m⊺​𝒌n{\boldsymbol{q}}_{m}^{\intercal}{\boldsymbol{k}}_{n} of Equation 2 as”
    From RoFormer (RoPE) · §Background and Related Work
  • Gopher2021 · cited 2×, 2 in Method
    “We use the autoregressive Transformer architecture detailed in Radford et al. 2019 with two modifications: we use RMSNorm (Zhang and Sennrich 2019) instead of LayerNorm (Ba et al. 2016), and we use the relative positiona…”
    From Gopher · §Method
  • GLaM2021 · cited 1×, 1 in Method
    “We replace the standard positional embedding with per-layer relative positional bias from Dai et al. 2019.”
    From GLaM · §Model Architecture
Abstract

Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling. We propose a novel neural architecture Transformer-XL that enables learning dependency beyond a fixed length without disrupting temporal coherence. It consists of a segment-level recurrence mechanism and a novel positional encoding scheme. Our method not only enables capturing longer-term dependency, but also resolves the context fragmentation problem. As a result, Transformer-XL learns dependency that is 80% longer than RNNs and 450% longer than vanilla Transformers, achieves better performance on both short and long sequences, and is up to 1,800+ times faster than vanilla Transformers during evaluation. Notably, we improve the state-of-the-art results of bpc/perplexity to 0.99 on enwiki8, 1.08 on text8, 18.3 on WikiText-103, 21.8 on One Billion Word, and 54.5 on Penn Treebank (without finetuning). When trained only on WikiText-103, Transformer-XL manages to generate reasonably coherent, novel text articles with thousands of tokens. Our code, pretrained models, and hyperparameters are available in both Tensorflow and PyTorch.