Paper Lineage
Esc
MethodMay 2019arXiv 1905.03197cs.CL

Unified Language Model Pre-training for Natural Language Understanding and Generation

Li Dong, Nan Yang, Wenhui Wang and 6 others

This paper presents a new Unified pre-trained Language Model (UniLM) that can be fine-tuned for both natural language understanding and generation tasks. The model is pre-trained using three types of language modeling tasks: unidirectional, bidirectional, and sequence-to-sequence prediction.

From the abstract

Built on

7 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (7)

  • BERT2018 · cited 9×, 5 in Method
    “The input representation follows that of BERT [9].”
    From this paper · §Unified Language Model Pre-training
  • Transformer2017 · cited 3×, 1 in Method
    “As shown in Figure 1, the pre-training optimizes the shared Transformer [43] network with respect to several unsupervised language modeling objectives, namely, unidirectional LM, bidirectional LM, and sequence-to-sequenc…”
    From this paper · §Unified Language Model Pre-training
  • BookCorpus (books & movies)2015 · cited 1×, 1 in Method
    “UniLM is initialized by BERTLARGE, and then pre-trained using English Wikipedia11 1 Wikipedia version: enwiki-20181101. and BookCorpus [53], which have been processed in the same way as [9].”
    From this paper · §Unified Language Model Pre-training
  • GNMT2016 · cited 1×, 1 in Method
    “Texts are tokenized to subword units by WordPiece [48].”
    From this paper · §Unified Language Model Pre-training
Show 3 more
  • CoQA2018 · cited 5×
    “We conduct experiments on the Stanford Question Answering Dataset (SQuAD) 2.0 [34], and Conversational Question Answering (CoQA) [35] datasets.”
    From this paper · §Experiments
  • SQuAD2016 · cited 3×
    “The task is to answer a question given a passage [33, 34, 15].”
    From this paper · §Experiments
  • ELMo2018 · cited 2×
    “Language model (LM) pre-training has substantially advanced the state of the art across a variety of natural language processing tasks [8, 29, 19, 31, 9, 1].”
    From this paper · §Introduction

Led to

  • Unified VLP2019 · cited 4×, 1 in Method
    “We follow the same scheme and consider two specific objectives: the bidirectional objective (bidirectional) as in BERT and the sequence to sequence objective (seq2seq), inspired by [\citeauthoryearDong et al.2019].”
    From Unified VLP · §Vision-Language Pre-training
  • T52019 · cited 5×
    “This approach has recently been used to obtain state-of-the-art results in many of the most common NLP benchmarks (Devlin et al. 2018; Yang et al. 2019; Dong et al. 2019; Liu et al. 2019c; Lan et al. 2019).”
    From T5 · §Introduction
  • BART2019 · cited 3×, 1 in Method
    “As in UniLM (Dong et al. 2019), we train a Masked Language Model with additional self-attention masks.”
    From BART · §Comparing Pre-training Objectives
  • VLMo2021 · cited 2×
    “Pre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”
    From VLMo · §Related Work
  • BEiT v22022 · cited 2×
    “The MIM method has achieved great success in language task (Devlin et al. 2019; Dong et al. 2019; Bao et al. 2020).”
    From BEiT v2 · §Related Work
  • BEiT-32022 · cited 2×
    “Following UniLM [17] and s2s-ft [5], BEiT-3 is used as a conditional generation model via masked finetuning.”
    From BEiT-3 · §Experiments on Vision and Vision-Language Tasks
  • BLIP-22023 · cited 1×, 1 in Method
    “We employ a multimodal causal self-attention mask to control query-text interaction, similar to the one used in UniLM (Dong et al. 2019).”
    From BLIP-2 · §Method
Abstract

This paper presents a new Unified pre-trained Language Model (UniLM) that can be fine-tuned for both natural language understanding and generation tasks. The model is pre-trained using three types of language modeling tasks: unidirectional, bidirectional, and sequence-to-sequence prediction. The unified modeling is achieved by employing a shared Transformer network and utilizing specific self-attention masks to control what context the prediction conditions on. UniLM compares favorably with BERT on the GLUE benchmark, and the SQuAD 2.0 and CoQA question answering tasks. Moreover, UniLM achieves new state-of-the-art results on five natural language generation datasets, including improving the CNN/DailyMail abstractive summarization ROUGE-L to 40.51 (2.04 absolute improvement), the Gigaword abstractive summarization ROUGE-L to 35.75 (0.86 absolute improvement), the CoQA generative question answering F1 score to 82.5 (37.1 absolute improvement), the SQuAD question generation BLEU-4 to 22.12 (3.75 absolute improvement), and the DSTC7 document-grounded dialog response generation NIST-4 to 2.67 (human performance is 2.65). The code and pre-trained models are available at https://github.com/microsoft/unilm.