Paper Lineage
Esc
MethodOct 2018arXiv 1810.04805cs.CL

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova

We introduce a new language representation model called BERT, which stands for Bidirectional Encoder Representations from Transformers. Unlike recent language representation models, BERT is designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers.

From the abstract

Built on

3 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (3)

  • ELMo2018 · cited 7×
    “ELMo and its predecessor Peters et al. 2017; Peters et al. 2018a generalize traditional word embedding research along a different dimension.”
    From this paper · §Related Work
  • Transformer2017 · cited 4×
    “BERT’s model architecture is a multi-layer bidirectional Transformer encoder based on the original implementation described in Vaswani et al. 2017 and released in the tensor2tensor library.11 1 https://github.com/tensorf…”
    From this paper · §BERT
  • SQuAD2016 · cited 3×
    “When integrating contextual word embeddings with existing task-specific architectures, ELMo advances the state of the art for several major NLP benchmarks Peters et al. 2018a including question answering Rajpurkar et al.…”
    From this paper · §Related Work

Led to

  • GPipe2018 · cited 2×
    “A similar phenomenon can also be observed in the context of natural language processing (Figure 1b) where simple shallow models of sentence representations [1, 2] are outperformed by their deeper and larger counterparts…”
    From GPipe · §Introduction
  • Transformer-XL2019 · cited 2×, 1 in Method
    “Second, though it is possible to use padding to respect the sentence or other semantic boundaries, in practice it has been standard practice to simply chunk long text into fixed-length segments due to improved efficiency…”
    From Transformer-XL · §Model
  • Adapters2019 · cited 7×
    “To perform classification with BERT, we follow the approach in Devlin et al. 2018.”
    From Adapters · §Experiments
  • UniLM2019 · cited 9×, 5 in Method
    “The input representation follows that of BERT [9].”
    From UniLM · §Unified Language Model Pre-training
  • BoolQ2019 · cited 4×, 1 in Method
    “Unsupervised: It is well known that unsupervised pre-training using language-modeling objectives Peters et al. 2018; Devlin et al. 2018; Radford et al. 2018, can improve performance on many tasks.”
    From BoolQ · §Training Yes/No QA Models
  • AMDIM2019 · cited 2×
    “Self-supervised learning is gaining popularity across the NLP, vision, and robotics communities – e.g., (Devlin et al. 2019; Logeswaran and Lee 2018; Sermanet et al. 2017; Dwibedi et al. 2018).”
    From AMDIM · §Related Work
  • XLNet2019 · cited 7×, 2 in Method
    “Replacing [MASK] with original tokens as in [10] does not solve the problem because original tokens can be only used with a small probability — otherwise Eq. (2) will be trivial to optimize.”
    From XLNet · §Proposed Method
  • RoBERTa2019 · cited 21×, 4 in Method
    “Unlike Devlin et al. 2019, we do not randomly inject short sequences, and we do not train with a reduced sequence length for the first 90% of updates.”
    From RoBERTa · §Experimental Setup
  • ViLBERT2019 · cited 9×, 1 in Method
    “In analogy to the training tasks in [12], we train our model on Conceptual Captions on two proxy tasks: predicting the semantics of masked words and image regions given the unmasked inputs, and predicting whether an imag…”
    From ViLBERT · §Introduction
  • VisualBERT2019 · cited 5×, 1 in Method
    “Our work is inspired by BERT (Devlin et al. 2019), a Transformer-based representation model for natural language.”
    From VisualBERT · §Related Work
  • Unicoder-VL2019 · cited 3×, 1 in Method
    “BERT [\citeauthoryearDevlin et al.2018] is a pre-trained model based on multi-layer Transformer [\citeauthoryearVaswani et al.2017].”
    From Unicoder-VL · §Approach
  • LXMERT2019 · cited 10×, 7 in Method
    “For the cross-modality output, following the practice in Devlin et al. 2019, we append a special token [CLS] (denoted as the top yellow block in the bottom branch of Fig. 1) before the sentence words, and the correspondi…”
    From LXMERT · §Model Architecture
  • VL-BERT2019 · cited 6×
    “After that, a serious of approaches are proposed for pre-training the generic representation, mainly based on Transformers, such as GPT (Radford et al. 2018), BERT (Devlin et al. 2018), GPT-2 (Radford et al. 2019), XLNet…”
    From VL-BERT · §Related Work
  • Megatron-LM2019 · cited 10×, 6 in Method
    “To analyze the effect of model size scaling on accuracy, we train both left-to-right GPT-2 (Radford et al. 2019) language models as well as BERT (Devlin et al. 2018) bidirectional transformers and evaluate them on severa…”
    From Megatron-LM · §Introduction
  • Unified VLP2019 · cited 7×, 1 in Method
    “Inspired by the recent success of pre-trained language models such as BERT [\citeauthoryearDevlin et al.2018] and GPT [\citeauthoryearRadford et al.2018, \citeauthoryearRadford et al.2019], there is a growing interest in…”
    From Unified VLP · §Introduction
  • FreeLB2019 · cited 2×
    “For BERT-base, we use the HuggingFace implementation44 4 https://github.com/huggingface/pytorch-transformers, and follow the single-task finetuning procedure as in Devlin et al. 2019.”
    From FreeLB · §Experiments
  • ALBERT2019 · cited 14×
    “BERT (Devlin et al. 2019) uses a loss based on predicting whether the second segment in a pair has been swapped with a segment from another document.”
    From ALBERT · §Related work
  • T52019 · cited 19×, 3 in Method
    “The Transformer was initially shown to be effective for machine translation, but it has subsequently been used in a wide variety of NLP settings (Radford et al. 2018; Devlin et al. 2018; McCann et al. 2018; Yu et al. 201…”
    From T5 · §Setup
  • BART2019 · cited 6×, 4 in Method
    “Following BERT Devlin et al. 2019, random tokens are sampled and replaced with [MASK] elements.”
    From BART · §Model
  • MoCo2019 · cited 2×
    “Beyond the simple instance discrimination task Wu2018a, it is possible to adopt MoCo for pretext tasks like masked auto-encoding, e.g., in language Devlin2019 and in vision Oord2018.”
    From MoCo · §Discussion and Conclusion
  • LPAQA (What LMs know)2019 · cited 4×
    “As for the models to probe, in our main experiments we use the standard BERT-base and BERT-large models (Devlin et al. 2019).”
    From LPAQA (What LMs know) · §Main Experiments
  • 12-in-12019 · cited 3×, 1 in Method
    “Internally, ViLBERT consists of two parallel BERT-style devlin2018bert models operating over image regions and text segments.”
    From 12-in-1 · §Approach
  • Meshed-Memory Transformer2019 · cited 2×
    “The recent advent of fully-attentive models, in which the recurrent relation is abandoned in favour of the use of self-attention, offers unique opportunities in terms of set and sequence modeling performances, as testifi…”
    From Meshed-Memory Transformer · §Introduction
  • Knowledge in LM parameters2020 · cited 4×, 2 in Method
    “Big, deep neural language models that have been pre-trained on unlabeled text have proven to be extremely performant when fine-tuned on downstream Natural Language Processing (NLP) tasks Devlin et al. 2018; Yang et al. 2…”
    From Knowledge in LM parameters · §Introduction
  • GPT-32020 · cited 4×
    “Work in this vein has successively increased model size: 213 million parameters [134] in the original paper, 300 million parameters [20], 1.5 billion parameters [117], 8 billion parameters [125], 11 billion parameters [1…”
    From GPT-3 · §Related Work
  • GShard2020 · cited 5×
    “For years, the fields have been continuously reporting new state of the art results using varieties of model architectures for computer vision tasks [57, 58, 7], for natural language understanding tasks [59, 60, 61], for…”
    From GShard · §Related Work
  • ConVIRT2020 · cited 1×, 1 in Method
    “For the text encoder fuf_{u}, we use a BERT encoder Devlin et al. 2019 followed by a max-pooling layer over all output vectors.”
    From ConVIRT · §Methods
  • ViT2020 · cited 4×
    “However, much of their success stems not only from their excellent scalability but also from large scale self-supervised pre-training (Devlin et al. 2019; Radford et al. 2018).”
    From ViT · §Experiments
  • VL-BERT meta-analysis2020 · cited 3×
    “In pursuit of this goal, many pretrained V&L models have been proposed in the last year, inspired by the success of pretraining in both computer vision (Sharif Razavian et al. 2014) and natural language processing (Devli…”
    From VL-BERT meta-analysis · §Introduction
  • DeiT2020 · cited 2×
    “Motivated by the success of attention-based models in Natural Language Processing [14, 52], there has been increasing interest in architectures leveraging attention mechanisms within convnets [2, 34, 61].”
    From DeiT · §Introduction
  • The Pile2020 · cited 2×
    “More recently, language models (Radford et al. 2018; Radford et al. 2019; Brown et al. 2020; Rosset 2019; Shoeybi et al. 2019) and masked language models (Devlin et al. 2019; Liu et al. 2019; Raffel et al. 2019) have bee…”
    From The Pile · §Related Work
  • Prefix-Tuning2021 · cited 2×
    “For extractive and abstractive summarization, researchers fine-tune masked language models (Devlin et al. 2019, e.g., BERT;) and encode-decoder models (Lewis et al. 2020, e.g., BART;) respectively Zhong et al. 2020; Liu…”
    From Prefix-Tuning · §Related Work
  • VL-T52021 · cited 4×, 2 in Method
    “RoI features and bounding box coordinates are encoded with a linear layer, while image ids and region ids are encoded with learned embeddings (Devlin et al. 2019).”
    From VL-T5 · §Model
  • ViLT2021 · cited 4×
    “Following the heuristics of Devlin et al. 2019, we randomly mask tt with the probability of 0.15.”
    From ViLT · §Vision-and-Language Transformer
  • Conceptual 12M2021 · cited 5×, 1 in Method
    “For instance, Zhou et al. [88] adapt BERT [25] to generate text.”
    From Conceptual 12M · §Evaluating Vision-and-Language Pre-Training Data
  • RoFormer (RoPE)2021 · cited 7×, 1 in Method
    “Previous work Devlin et al. 2019; Lan et al. 2020; Clark et al. 2020; Radford et al. 2019; Radford and Narasimhan 2018 introduced the use of a set of trainable vectors 𝒑i∈{𝒑t}t=1L{\boldsymbol{p}}_{i}\in\{{\boldsymbol{p…”
    From RoFormer (RoPE) · §Background and Related Work
  • BEiT2021 · cited 3×, 2 in Method
    “The image patches {𝒙ip}i=1N\{{\bm{x}}^{p}_{i}\}_{i=1}^{N} are flattened into vectors and are linearly projected, which is similar to word embeddings in BERT [13].”
    From BEiT · §Methods
  • Codex2021 · cited 3×
    “Following the success of large natural language models (Devlin et al. 2018; Radford et al. 2019; Liu et al. 2019; Raffel et al. 2020; Brown et al. 2020) large scale Transformers have also been applied towards program syn…”
    From Codex · §Related Work
  • ALBEF2021 · cited 1×, 1 in Method
    “The text encoder is initialized using the first 6 layers of the BERTbase [40] model, and the multimodal encoder is initialized using the last 6 layers of the BERTbase.”
    From ALBEF · §ALBEF Pre-training
  • SimVLM2021 · cited 5×
    “Self-supervised textual representation learning (Devlin et al. 2018; Radford et al. 2018; Radford et al. 2019; Liu et al. 2019; Yang et al. 2019; Raffel et al. 2019; Brown et al. 2020) based on Transformers (Vaswani et a…”
    From SimVLM · §Introduction
  • ViT-VQGAN2021 · cited 2×
    “Driven by the effectiveness of VQVAE and progress in sequence modeling (Vaswani et al. 2017; Devlin et al. 2019), many approaches follow the two-stage paradigm.”
    From ViT-VQGAN · §Related Work
  • VLMo2021 · cited 6×, 3 in Method
    “Following BERT [10], we tokenize the text to subword units by WordPiece [47].”
    From VLMo · §Methods
  • METER2021 · cited 3×, 3 in Method
    “Following BERT devlin2018bert and RoBERTa liu2019roberta, VLP models tan-bansal-2019-lxmert; li2019visualbert; lu2019vilbert; su2019vl; chen2020uniter; li2020oscar first segment the input sentence into a sequence of subw…”
    From METER · §The Meter Framework
  • MAE2021 · cited 9×, 2 in Method
    “The solutions, based on autoregressive language modeling in GPT Radford2018; Radford2019; Brown2020 and masked autoencoding in BERT Devlin2019, are conceptually simple: they remove a portion of the data and learn to pred…”
    From MAE · §Introduction
  • Swin V22021 · cited 3×
    “It significantly improves a model’s performance on language tasks devlin2018bert; radford2019language; raffel2019t5; Turing-17B; fedus2021switch; Megatron-Turing-530B and the model demonstrates amazing few-shot capabilit…”
    From Swin V2 · §Introduction
  • Gopher2021 · cited 2×, 2 in Method
    “Whilst there are other objectives towards modelling a sequence, such as modelling masked tokens given bi-directional context (Mikolov et al. 2013; Devlin et al. 2019) and modelling all permutations of the sequence (Yang…”
    From Gopher · §Background
  • ViTCAP2021 · cited 2×
    “The transformer architecture and its instantiations (e.g., BERT devlin2018bert, GPT brown2020language) are well-known for their remarkable performances on natural language processing tasks, which are mostly attributed to…”
    From ViTCAP · §ViTCAP
  • GLaM2021 · cited 2×
    “More recently, models that used Transformers (Vaswani et al. 2017) showed that larger models with self-supervision on unlabeled data could yield significant improvements on NLP tasks (Devlin et al. 2019; Yang et al. 2019…”
    From GLaM · §Related Work
  • Fairseq MoE LMs2021 · cited 4×, 3 in Method
    “This is in contrast to the traditional approach of augmenting LMs with task-specific heads, followed by supervised fine-tuning (Devlin et al. 2019; Raffel et al. 2020).”
    From Fairseq MoE LMs · §Background and Related Work
  • Megatron-Turing NLG2022 · cited 2×
    “This trend is continued when large-scale pretraining with transformer architectures becomes popular, with BERT [12] scaling up to 300 million parameters, followed by GPT-2 [52] at 1.5 billion parameters.”
    From Megatron-Turing NLG · §Related Works
  • BLIP2022 · cited 2×, 1 in Method
    “The text encoder is the same as BERT (Devlin et al. 2019), where a [CLS] token is appended to the beginning of the text input to summarize the sentence.”
    From BLIP · §Method
  • PaLM2022 · cited 3×
    “Broadly, language modeling refers to approaches for predicting either the next token in a sequence or for predicting masked spans (Devlin et al. 2019; Raffel et al. 2020).”
    From PaLM · §Related Work
  • Flamingo2022 · cited 2×
    “The paradigm of first pretraining on a vast amount of data followed by an adaptation on a downstream task has become standard [75, 32, 52, 44, 23, 87, 108, 11].”
    From Flamingo · §Related work
  • CoCa2022 · cited 2×
    “Pretraining ConvNets [18] or Transformers [19] on large-scale annotated data such as ImageNet [6, 7, 8], Instagram [20] or JFT [21] has become a popular strategy towards solving visual recognition problems including clas…”
    From CoCa · §Related Work
  • VL-BEiT2022 · cited 3×, 1 in Method
    “Following BERT [7], we randomly mask 15% tokens of monomodal text data.”
    From VL-BEiT · §Methods
  • BIG-bench2022 · cited 4×
    “Task cites: (Devlin et al. 2018; Wu et al. 2016; Won et al. 2021; Xu et al. 2018; Edmiston & Stratos 2018; El-Kishky et al. 2019)”
    From BIG-bench · §Author contributions
  • BEiT v22022 · cited 2×
    “The MIM method has achieved great success in language task (Devlin et al. 2019; Dong et al. 2019; Bao et al. 2020).”
    From BEiT v2 · §Related Work
  • BEiT-32022 · cited 3×
    “Second, the pretraining task based on masked data modeling has been successfully applied to various modalities, such as texts [14], images [3, 40], and image-text pairs [7].”
    From BEiT-3 · §Introduction: The Big Convergence
  • U-PaLM2022 · cited 2×
    “While there have been many paradigms and self-supervision methods proposed to train these models (Devlin et al. 2018; Clark et al. 2020b; Yang et al. 2019; Raffel et al. 2019), to this date most large language models (i.…”
    From U-PaLM · §Related Work
  • BLIP-22023 · cited 1×, 1 in Method
    “We initialize Q-Former with the pre-trained weights of BERTbase (Devlin et al. 2019), whereas the cross-attention layers are randomly initialized.”
    From BLIP-2 · §Method
Abstract

We introduce a new language representation model called BERT, which stands for Bidirectional Encoder Representations from Transformers. Unlike recent language representation models, BERT is designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers. As a result, the pre-trained BERT model can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks, such as question answering and language inference, without substantial task-specific architecture modifications. BERT is conceptually simple and empirically powerful. It obtains new state-of-the-art results on eleven natural language processing tasks, including pushing the GLUE score to 80.5% (7.7% point absolute improvement), MultiNLI accuracy to 86.7% (4.6% absolute improvement), SQuAD v1.1 question answering Test F1 to 93.2 (1.5 point absolute improvement) and SQuAD v2.0 Test F1 to 83.1 (5.1 point absolute improvement).