Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling. We propose a novel neural architecture Transformer-XL that enables learning dependency beyond a fixed length without disrupting temporal coherence.
Also cited · not yet reviewed (3)
- Transformer2017 · cited 5×, 3 in Method“Firstly, in the standard Transformer (Vaswani et al. 2017), the attention score between query qiq_{i} and key vector kjk_{j} within the same segment can be decomposed as”From this paper · §Model
- ELMo2018 · cited 2×, 1 in Method“Second, though it is possible to use padding to respect the sentence or other semantic boundaries, in practice it has been standard practice to simply chunk long text into fixed-length segments due to improved efficiency…”From this paper · §Model
- BERT2018 · cited 2×, 1 in Method“Second, though it is possible to use padding to respect the sentence or other semantic boundaries, in practice it has been standard practice to simply chunk long text into fixed-length segments due to improved efficiency…”From this paper · §Model
Led to
- XLNet2019 · cited 4×, 2 in Method“Inspired by the latest advancements in AR language modeling, XLNet integrates the segment recurrence mechanism and relative encoding scheme of Transformer-XL [9] into pretraining, which empirically improves the performan…”From XLNet · §Introduction
- Megatron-LM2019 · cited 2×, 1 in Method“Recent parallel work (Ramachandran et al. 2016; Howard & Ruder 2018; Radford et al. 2018; Devlin et al. 2018; Liu et al. 2019b; Dai et al. 2019; Yang et al. 2019; Liu et al. 2019a; Lan et al. 2019) further builds upon th…”From Megatron-LM · §Background and Challenges
- RoFormer (RoPE)2021 · cited 2×, 1 in Method“Keeping the form of Equation 3, the authors Dai et al. 2019 have proposed to decompose 𝒒m⊺𝒌n{\boldsymbol{q}}_{m}^{\intercal}{\boldsymbol{k}}_{n} of Equation 2 as”From RoFormer (RoPE) · §Background and Related Work
- Gopher2021 · cited 2×, 2 in Method“We use the autoregressive Transformer architecture detailed in Radford et al. 2019 with two modifications: we use RMSNorm (Zhang and Sennrich 2019) instead of LayerNorm (Ba et al. 2016), and we use the relative positiona…”From Gopher · §Method
- GLaM2021 · cited 1×, 1 in Method“We replace the standard positional embedding with per-layer relative positional bias from Dai et al. 2019.”From GLaM · §Model Architecture
Abstract
Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling. We propose a novel neural architecture Transformer-XL that enables learning dependency beyond a fixed length without disrupting temporal coherence. It consists of a segment-level recurrence mechanism and a novel positional encoding scheme. Our method not only enables capturing longer-term dependency, but also resolves the context fragmentation problem. As a result, Transformer-XL learns dependency that is 80% longer than RNNs and 450% longer than vanilla Transformers, achieves better performance on both short and long sequences, and is up to 1,800+ times faster than vanilla Transformers during evaluation. Notably, we improve the state-of-the-art results of bpc/perplexity to 0.99 on enwiki8, 1.08 on text8, 18.3 on WikiText-103, 21.8 on One Billion Word, and 54.5 on Penn Treebank (without finetuning). When trained only on WikiText-103, Transformer-XL manages to generate reasonably coherent, novel text articles with thousands of tokens. Our code, pretrained models, and hyperparameters are available in both Tensorflow and PyTorch.