Paper Lineage
Esc
MethodJul 2019arXiv 1907.11692cs.CL

RoBERTa: A Robustly Optimized BERT Pretraining Approach

Yinhan Liu, Myle Ott, Naman Goyal and 7 others

Language model pretraining has led to significant performance gains but careful comparison between different approaches is challenging. Training is computationally expensive, often done on private datasets of different sizes, and, as we will show, hyperparameter choices have significant impact on the final results.

From the abstract

Built on

6 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (6)

  • BERT2018 · cited 21×, 4 in Method
    “Unlike Devlin et al. 2019, we do not randomly inject short sequences, and we do not train with a reduced sequence length for the first 90% of updates.”
    From this paper · §Experimental Setup
  • XLNet2019 · cited 13×, 2 in Method
    “Several efforts have trained on datasets larger and more diverse than the original BERT Radford et al. 2019; Yang et al. 2019; Zellers et al. 2019.”
    From this paper · §Experimental Setup
  • BookCorpus (books & movies)2015 · cited 2×, 2 in Method
    “BookCorpus Zhu et al. 2015 plus English Wikipedia.”
    From this paper · §Experimental Setup
  • SQuAD2016 · cited 2×, 2 in Method
    “The General Language Understanding Evaluation (GLUE) benchmark Wang et al. 2019b is a collection of 9 datasets for evaluating natural language understanding systems.66 6 The datasets are: CoLA Warstadt et al. 2018, Stanf…”
    From this paper · §Experimental Setup
Show 2 more
  • Transformer2017 · cited 1×, 1 in Method
    “BERT uses the now ubiquitous transformer architecture Vaswani et al. 2017, which we will not review in detail.”
    From this paper · §Background
  • ELMo2018 · cited 2×
    “Pretraining methods have been designed with different training objectives, including language modeling Dai and Le 2015; Peters et al. 2018; Howard and Ruder 2018, machine translation McCann et al. 2017, and masked langua…”
    From this paper · §Related Work

Led to

  • Unicoder-VL2019 · cited 2×
    “Latest pre-trained NLP models are based on multi-layer Transformer, such as GPT [\citeauthoryearRadford et al.2018], BERT [\citeauthoryearDevlin et al.2018], XLNet (Yang et al., 2019) and RoBERTa[\citeauthoryearLiu et al…”
    From Unicoder-VL · §Related Work
  • Megatron-LM2019 · cited 3×, 1 in Method
    “For finetuning, we follow the same procedure as (Liu et al. 2019b).”
    From Megatron-LM · §Experiments
  • Unified VLP2019 · cited 1×, 1 in Method
    “This coincidentally agrees with a concurrent work of RoBERTa [\citeauthoryearLiu et al.2019b].”
    From Unified VLP · §Vision-Language Pre-training
  • FreeLB2019 · cited 4×
    “CommonsenseQA Similar to the training strategy in Liu et al. 2019b, we construct five inputs for each question by concatenating the question and each answer separately, then encode each input with the representation of t…”
    From FreeLB · §Experiments
  • ALBERT2019 · cited 13×
    “One of the most compelling signs of these breakthroughs is the evolution of machine performance on a reading comprehension task designed for middle and high-school English exams in China, the RACE test (Lai et al. 2017):…”
    From ALBERT · §Introduction
  • T52019 · cited 10×, 2 in Method
    “However, we opted to create a new data set because prior data sets use a more limited set of filtering heuristics, are not publicly available, and/or are different in scope (e.g. are limited to News data (Zellers et al.…”
    From T5 · §Setup
  • BART2019 · cited 7×, 3 in Method
    “Recent work has shown that downstream performance can dramatically improve when pre-training is scaled to large batch sizes (Yang et al. 2019; Liu et al. 2019) and corpora.”
    From BART · §Large-scale Pre-training Experiments
  • “Big, deep neural language models that have been pre-trained on unlabeled text have proven to be extremely performant when fine-tuned on downstream Natural Language Processing (NLP) tasks Devlin et al. 2018; Yang et al. 2…”
    From Knowledge in LM parameters · §Introduction
  • GPT-32020 · cited 2×
    “This last paradigm has led to substantial progress on many challenging NLP tasks such as reading comprehension, question answering, textual entailment, and many others, and has continued to advance based on new architect…”
    From GPT-3 · §Introduction
  • VirTex2020 · cited 2×, 1 in Method
    “However, following BERT [64], many large-scale models [82, 83] instead use masked language models (MLMs): some tokens are randomly masked and are predicted by the model.”
    From VirTex · §Method
  • The Pile2020 · cited 2×
    “More recently, language models (Radford et al. 2018; Radford et al. 2019; Brown et al. 2020; Rosset 2019; Shoeybi et al. 2019) and masked language models (Devlin et al. 2019; Liu et al. 2019; Raffel et al. 2019) have bee…”
    From The Pile · §Related Work
  • METER2021 · cited 5×, 3 in Method
    “Specifically, as shown in Figure 1, we dissect the model designs along multiple dimensions, including vision encoders (e.g., CLIP-ViT radford2021learning, Swin transformer liu2021swin), text encoders (e.g., RoBERTa liu20…”
    From METER · §Introduction
  • Florence2021 · cited 1×, 1 in Method
    “In the Florence V+L adaptation model, we replace the image encoder of METER (Dou et al. 2021) with Florence pretrained model CoSwin, and use a pretrained Roberta (Liu et al. 2019) as the language encoder, shown in Figure…”
    From Florence · §Approach
  • Fairseq MoE LMs2021 · cited 5×, 4 in Method
    “We pretrain our models on a union of six English-language datasets, including the five datasets used to pretrain RoBERTa (Liu et al. 2019) and the English subset of CC100, totalling 112B tokens corresponding to 453GB:”
    From Fairseq MoE LMs · §Experimental Setup
  • OPT2022 · cited 3×, 2 in Method
    “The pre-training corpus contains a concatenation of datasets used in RoBERTa Liu et al. 2019b, the Pile Gao et al. 2021a, and PushShift.io Reddit Baumgartner et al. 2020; Roller et al. 2021.”
    From OPT · §Method
  • BIG-bench2022 · cited 2×
    “Task cites: (Brown et al. 2020; Devlin et al. 2018; Lan et al. 2019; Liu et al. 2019; Radford et al. 2019; Annamoradnejad & Zoghi 2020; Chen & Soo 2018; Mao & Liu 2019; Weller & Seppi 2019; Khodak et al. 2017; Ghosh et a…”
    From BIG-bench · §Author contributions
  • BEiT-32022 · cited 1×, 1 in Method
    “For monomodal data, we use 1414M images from ImageNet-21K and 160160GB text corpora [4] from English Wikipedia, BookCorpus [64], OpenWebText22 2 http://skylion007.github.io/OpenWebTextCorpus, CC-News [33], and Stories [5…”
    From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
Abstract

Language model pretraining has led to significant performance gains but careful comparison between different approaches is challenging. Training is computationally expensive, often done on private datasets of different sizes, and, as we will show, hyperparameter choices have significant impact on the final results. We present a replication study of BERT pretraining (Devlin et al., 2019) that carefully measures the impact of many key hyperparameters and training data size. We find that BERT was significantly undertrained, and can match or exceed the performance of every model published after it. Our best model achieves state-of-the-art results on GLUE, RACE and SQuAD. These results highlight the importance of previously overlooked design choices, and raise questions about the source of recently reported improvements. We release our models and code.