Paper Lineage
Esc
MethodJun 2019arXiv 1906.08237cs.CL

XLNet: Generalized Autoregressive Pretraining for Language Understanding

Zhilin Yang, Zihang Dai, Yiming Yang and 3 others

With the capability of modeling bidirectional contexts, denoising autoencoding based pretraining like BERT achieves better performance than pretraining approaches based on autoregressive language modeling. However, relying on corrupting the input with masks, BERT neglects dependency between the masked positions and suffers from a pretrain-finetune discrepancy.

From the abstract

Built on

4 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (4)

  • BERT2018 · cited 7×, 2 in Method
    “Replacing [MASK] with original tokens as in [10] does not solve the problem because original tokens can be only used with a small probability — otherwise Eq. (2) will be trivial to optimize.”
    From this paper · §Proposed Method
  • Transformer-XL2019 · cited 4×, 2 in Method
    “Inspired by the latest advancements in AR language modeling, XLNet integrates the segment recurrence mechanism and relative encoding scheme of Transformer-XL [9] into pretraining, which empirically improves the performan…”
    From this paper · §Introduction
  • Transformer2017 · cited 1×, 1 in Method
    “where Q, K, V denote the query, key, and value in an attention operation [33].”
    From this paper · §Proposed Method
  • ELMo2018 · cited 2×
    “Unsupervised representation learning has been highly successful in the domain of natural language processing [7, 22, 27, 28, 10].”
    From this paper · §Introduction

Led to

  • RoBERTa2019 · cited 13×, 2 in Method
    “Several efforts have trained on datasets larger and more diverse than the original BERT Radford et al. 2019; Yang et al. 2019; Zellers et al. 2019.”
    From RoBERTa · §Experimental Setup
  • Megatron-LM2019 · cited 2×, 1 in Method
    “Recent parallel work (Ramachandran et al. 2016; Howard & Ruder 2018; Radford et al. 2018; Devlin et al. 2018; Liu et al. 2019b; Dai et al. 2019; Yang et al. 2019; Liu et al. 2019a; Lan et al. 2019) further builds upon th…”
    From Megatron-LM · §Background and Challenges
  • ALBERT2019 · cited 9×
    “Like BERT, we use a vocabulary size of 30,000, tokenized using SentencePiece (Kudo & Richardson 2018) as in XLNet (Yang et al. 2019).”
    From ALBERT · §Experimental Results
  • T52019 · cited 10×, 1 in Method
    “It has recently also become common to use models consisting of a single Transformer layer stack, with varying forms of self-attention used to produce architectures appropriate for language modeling (Radford et al. 2018;…”
    From T5 · §Setup
  • BART2019 · cited 7×, 4 in Method
    “Based on XLNet (Yang et al. 2019), we sample 1/6 of the tokens, and generate them in a random order autoregressively.”
    From BART · §Comparing Pre-training Objectives
  • Knowledge in LM parameters2020 · cited 2×, 1 in Method
    “Big, deep neural language models that have been pre-trained on unlabeled text have proven to be extremely performant when fine-tuned on downstream Natural Language Processing (NLP) tasks Devlin et al. 2018; Yang et al. 2…”
    From Knowledge in LM parameters · §Introduction
  • GPT-32020 · cited 2×
    “This last paradigm has led to substantial progress on many challenging NLP tasks such as reading comprehension, question answering, textual entailment, and many others, and has continued to advance based on new architect…”
    From GPT-3 · §Introduction
  • Switch Transformer2021 · cited 1×, 1 in Method
    “On ANLI (Nie et al. 2019), Switch XXL improves over the prior state-of-the-art to get a 65.7 accuracy versus the prior best of 49.4 (Yang et al. 2020).”
    From Switch Transformer · §Designing Models with Data, Model, and Expert-Parallelism
  • Gopher2021 · cited 1×, 1 in Method
    “Whilst there are other objectives towards modelling a sequence, such as modelling masked tokens given bi-directional context (Mikolov et al. 2013; Devlin et al. 2019) and modelling all permutations of the sequence (Yang…”
    From Gopher · §Background
Abstract

With the capability of modeling bidirectional contexts, denoising autoencoding based pretraining like BERT achieves better performance than pretraining approaches based on autoregressive language modeling. However, relying on corrupting the input with masks, BERT neglects dependency between the masked positions and suffers from a pretrain-finetune discrepancy. In light of these pros and cons, we propose XLNet, a generalized autoregressive pretraining method that (1) enables learning bidirectional contexts by maximizing the expected likelihood over all permutations of the factorization order and (2) overcomes the limitations of BERT thanks to its autoregressive formulation. Furthermore, XLNet integrates ideas from Transformer-XL, the state-of-the-art autoregressive model, into pretraining. Empirically, under comparable experiment settings, XLNet outperforms BERT on 20 tasks, often by a large margin, including question answering, natural language inference, sentiment analysis, and document ranking.