BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
We introduce a new language representation model called BERT, which stands for Bidirectional Encoder Representations from Transformers. Unlike recent language representation models, BERT is designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers.
Also cited · not yet reviewed (3)
- ELMo2018 · cited 7דELMo and its predecessor Peters et al. 2017; Peters et al. 2018a generalize traditional word embedding research along a different dimension.”From this paper · §Related Work
- Transformer2017 · cited 4דBERT’s model architecture is a multi-layer bidirectional Transformer encoder based on the original implementation described in Vaswani et al. 2017 and released in the tensor2tensor library.11 1 https://github.com/tensorf…”From this paper · §BERT
- SQuAD2016 · cited 3דWhen integrating contextual word embeddings with existing task-specific architectures, ELMo advances the state of the art for several major NLP benchmarks Peters et al. 2018a including question answering Rajpurkar et al.…”From this paper · §Related Work
Led to
- GPipe2018 · cited 2דA similar phenomenon can also be observed in the context of natural language processing (Figure 1b) where simple shallow models of sentence representations [1, 2] are outperformed by their deeper and larger counterparts…”From GPipe · §Introduction
- Transformer-XL2019 · cited 2×, 1 in Method“Second, though it is possible to use padding to respect the sentence or other semantic boundaries, in practice it has been standard practice to simply chunk long text into fixed-length segments due to improved efficiency…”From Transformer-XL · §Model
- Adapters2019 · cited 7דTo perform classification with BERT, we follow the approach in Devlin et al. 2018.”From Adapters · §Experiments
- UniLM2019 · cited 9×, 5 in Method“The input representation follows that of BERT [9].”From UniLM · §Unified Language Model Pre-training
- BoolQ2019 · cited 4×, 1 in Method“Unsupervised: It is well known that unsupervised pre-training using language-modeling objectives Peters et al. 2018; Devlin et al. 2018; Radford et al. 2018, can improve performance on many tasks.”From BoolQ · §Training Yes/No QA Models
- AMDIM2019 · cited 2דSelf-supervised learning is gaining popularity across the NLP, vision, and robotics communities – e.g., (Devlin et al. 2019; Logeswaran and Lee 2018; Sermanet et al. 2017; Dwibedi et al. 2018).”From AMDIM · §Related Work
- XLNet2019 · cited 7×, 2 in Method“Replacing [MASK] with original tokens as in [10] does not solve the problem because original tokens can be only used with a small probability — otherwise Eq. (2) will be trivial to optimize.”From XLNet · §Proposed Method
- RoBERTa2019 · cited 21×, 4 in Method“Unlike Devlin et al. 2019, we do not randomly inject short sequences, and we do not train with a reduced sequence length for the first 90% of updates.”From RoBERTa · §Experimental Setup
- ViLBERT2019 · cited 9×, 1 in Method“In analogy to the training tasks in [12], we train our model on Conceptual Captions on two proxy tasks: predicting the semantics of masked words and image regions given the unmasked inputs, and predicting whether an imag…”From ViLBERT · §Introduction
- VisualBERT2019 · cited 5×, 1 in Method“Our work is inspired by BERT (Devlin et al. 2019), a Transformer-based representation model for natural language.”From VisualBERT · §Related Work
- Unicoder-VL2019 · cited 3×, 1 in Method“BERT [\citeauthoryearDevlin et al.2018] is a pre-trained model based on multi-layer Transformer [\citeauthoryearVaswani et al.2017].”From Unicoder-VL · §Approach
- LXMERT2019 · cited 10×, 7 in Method“For the cross-modality output, following the practice in Devlin et al. 2019, we append a special token [CLS] (denoted as the top yellow block in the bottom branch of Fig. 1) before the sentence words, and the correspondi…”From LXMERT · §Model Architecture
- VL-BERT2019 · cited 6דAfter that, a serious of approaches are proposed for pre-training the generic representation, mainly based on Transformers, such as GPT (Radford et al. 2018), BERT (Devlin et al. 2018), GPT-2 (Radford et al. 2019), XLNet…”From VL-BERT · §Related Work
- Megatron-LM2019 · cited 10×, 6 in Method“To analyze the effect of model size scaling on accuracy, we train both left-to-right GPT-2 (Radford et al. 2019) language models as well as BERT (Devlin et al. 2018) bidirectional transformers and evaluate them on severa…”From Megatron-LM · §Introduction
- Unified VLP2019 · cited 7×, 1 in Method“Inspired by the recent success of pre-trained language models such as BERT [\citeauthoryearDevlin et al.2018] and GPT [\citeauthoryearRadford et al.2018, \citeauthoryearRadford et al.2019], there is a growing interest in…”From Unified VLP · §Introduction
- FreeLB2019 · cited 2דFor BERT-base, we use the HuggingFace implementation44 4 https://github.com/huggingface/pytorch-transformers, and follow the single-task finetuning procedure as in Devlin et al. 2019.”From FreeLB · §Experiments
- ALBERT2019 · cited 14דBERT (Devlin et al. 2019) uses a loss based on predicting whether the second segment in a pair has been swapped with a segment from another document.”From ALBERT · §Related work
- T52019 · cited 19×, 3 in Method“The Transformer was initially shown to be effective for machine translation, but it has subsequently been used in a wide variety of NLP settings (Radford et al. 2018; Devlin et al. 2018; McCann et al. 2018; Yu et al. 201…”From T5 · §Setup
- BART2019 · cited 6×, 4 in Method“Following BERT Devlin et al. 2019, random tokens are sampled and replaced with [MASK] elements.”From BART · §Model
- MoCo2019 · cited 2דBeyond the simple instance discrimination task Wu2018a, it is possible to adopt MoCo for pretext tasks like masked auto-encoding, e.g., in language Devlin2019 and in vision Oord2018.”From MoCo · §Discussion and Conclusion
- LPAQA (What LMs know)2019 · cited 4דAs for the models to probe, in our main experiments we use the standard BERT-base and BERT-large models (Devlin et al. 2019).”From LPAQA (What LMs know) · §Main Experiments
- 12-in-12019 · cited 3×, 1 in Method“Internally, ViLBERT consists of two parallel BERT-style devlin2018bert models operating over image regions and text segments.”From 12-in-1 · §Approach
- Meshed-Memory Transformer2019 · cited 2דThe recent advent of fully-attentive models, in which the recurrent relation is abandoned in favour of the use of self-attention, offers unique opportunities in terms of set and sequence modeling performances, as testifi…”From Meshed-Memory Transformer · §Introduction
- Knowledge in LM parameters2020 · cited 4×, 2 in Method“Big, deep neural language models that have been pre-trained on unlabeled text have proven to be extremely performant when fine-tuned on downstream Natural Language Processing (NLP) tasks Devlin et al. 2018; Yang et al. 2…”From Knowledge in LM parameters · §Introduction
- GPT-32020 · cited 4דWork in this vein has successively increased model size: 213 million parameters [134] in the original paper, 300 million parameters [20], 1.5 billion parameters [117], 8 billion parameters [125], 11 billion parameters [1…”From GPT-3 · §Related Work
- GShard2020 · cited 5דFor years, the fields have been continuously reporting new state of the art results using varieties of model architectures for computer vision tasks [57, 58, 7], for natural language understanding tasks [59, 60, 61], for…”From GShard · §Related Work
- ConVIRT2020 · cited 1×, 1 in Method“For the text encoder fuf_{u}, we use a BERT encoder Devlin et al. 2019 followed by a max-pooling layer over all output vectors.”From ConVIRT · §Methods
- ViT2020 · cited 4דHowever, much of their success stems not only from their excellent scalability but also from large scale self-supervised pre-training (Devlin et al. 2019; Radford et al. 2018).”From ViT · §Experiments
- VL-BERT meta-analysis2020 · cited 3דIn pursuit of this goal, many pretrained V&L models have been proposed in the last year, inspired by the success of pretraining in both computer vision (Sharif Razavian et al. 2014) and natural language processing (Devli…”From VL-BERT meta-analysis · §Introduction
- DeiT2020 · cited 2דMotivated by the success of attention-based models in Natural Language Processing [14, 52], there has been increasing interest in architectures leveraging attention mechanisms within convnets [2, 34, 61].”From DeiT · §Introduction
- The Pile2020 · cited 2דMore recently, language models (Radford et al. 2018; Radford et al. 2019; Brown et al. 2020; Rosset 2019; Shoeybi et al. 2019) and masked language models (Devlin et al. 2019; Liu et al. 2019; Raffel et al. 2019) have bee…”From The Pile · §Related Work
- Prefix-Tuning2021 · cited 2דFor extractive and abstractive summarization, researchers fine-tune masked language models (Devlin et al. 2019, e.g., BERT;) and encode-decoder models (Lewis et al. 2020, e.g., BART;) respectively Zhong et al. 2020; Liu…”From Prefix-Tuning · §Related Work
- VL-T52021 · cited 4×, 2 in Method“RoI features and bounding box coordinates are encoded with a linear layer, while image ids and region ids are encoded with learned embeddings (Devlin et al. 2019).”From VL-T5 · §Model
- ViLT2021 · cited 4דFollowing the heuristics of Devlin et al. 2019, we randomly mask tt with the probability of 0.15.”From ViLT · §Vision-and-Language Transformer
- Conceptual 12M2021 · cited 5×, 1 in Method“For instance, Zhou et al. [88] adapt BERT [25] to generate text.”From Conceptual 12M · §Evaluating Vision-and-Language Pre-Training Data
- RoFormer (RoPE)2021 · cited 7×, 1 in Method“Previous work Devlin et al. 2019; Lan et al. 2020; Clark et al. 2020; Radford et al. 2019; Radford and Narasimhan 2018 introduced the use of a set of trainable vectors 𝒑i∈{𝒑t}t=1L{\boldsymbol{p}}_{i}\in\{{\boldsymbol{p…”From RoFormer (RoPE) · §Background and Related Work
- BEiT2021 · cited 3×, 2 in Method“The image patches {𝒙ip}i=1N\{{\bm{x}}^{p}_{i}\}_{i=1}^{N} are flattened into vectors and are linearly projected, which is similar to word embeddings in BERT [13].”From BEiT · §Methods
- Codex2021 · cited 3דFollowing the success of large natural language models (Devlin et al. 2018; Radford et al. 2019; Liu et al. 2019; Raffel et al. 2020; Brown et al. 2020) large scale Transformers have also been applied towards program syn…”From Codex · §Related Work
- ALBEF2021 · cited 1×, 1 in Method“The text encoder is initialized using the first 6 layers of the BERTbase [40] model, and the multimodal encoder is initialized using the last 6 layers of the BERTbase.”From ALBEF · §ALBEF Pre-training
- SimVLM2021 · cited 5דSelf-supervised textual representation learning (Devlin et al. 2018; Radford et al. 2018; Radford et al. 2019; Liu et al. 2019; Yang et al. 2019; Raffel et al. 2019; Brown et al. 2020) based on Transformers (Vaswani et a…”From SimVLM · §Introduction
- ViT-VQGAN2021 · cited 2דDriven by the effectiveness of VQVAE and progress in sequence modeling (Vaswani et al. 2017; Devlin et al. 2019), many approaches follow the two-stage paradigm.”From ViT-VQGAN · §Related Work
- VLMo2021 · cited 6×, 3 in Method“Following BERT [10], we tokenize the text to subword units by WordPiece [47].”From VLMo · §Methods
- METER2021 · cited 3×, 3 in Method“Following BERT devlin2018bert and RoBERTa liu2019roberta, VLP models tan-bansal-2019-lxmert; li2019visualbert; lu2019vilbert; su2019vl; chen2020uniter; li2020oscar first segment the input sentence into a sequence of subw…”From METER · §The Meter Framework
- MAE2021 · cited 9×, 2 in Method“The solutions, based on autoregressive language modeling in GPT Radford2018; Radford2019; Brown2020 and masked autoencoding in BERT Devlin2019, are conceptually simple: they remove a portion of the data and learn to pred…”From MAE · §Introduction
- Swin V22021 · cited 3דIt significantly improves a model’s performance on language tasks devlin2018bert; radford2019language; raffel2019t5; Turing-17B; fedus2021switch; Megatron-Turing-530B and the model demonstrates amazing few-shot capabilit…”From Swin V2 · §Introduction
- Gopher2021 · cited 2×, 2 in Method“Whilst there are other objectives towards modelling a sequence, such as modelling masked tokens given bi-directional context (Mikolov et al. 2013; Devlin et al. 2019) and modelling all permutations of the sequence (Yang…”From Gopher · §Background
- ViTCAP2021 · cited 2דThe transformer architecture and its instantiations (e.g., BERT devlin2018bert, GPT brown2020language) are well-known for their remarkable performances on natural language processing tasks, which are mostly attributed to…”From ViTCAP · §ViTCAP
- GLaM2021 · cited 2דMore recently, models that used Transformers (Vaswani et al. 2017) showed that larger models with self-supervision on unlabeled data could yield significant improvements on NLP tasks (Devlin et al. 2019; Yang et al. 2019…”From GLaM · §Related Work
- Fairseq MoE LMs2021 · cited 4×, 3 in Method“This is in contrast to the traditional approach of augmenting LMs with task-specific heads, followed by supervised fine-tuning (Devlin et al. 2019; Raffel et al. 2020).”From Fairseq MoE LMs · §Background and Related Work
- Megatron-Turing NLG2022 · cited 2דThis trend is continued when large-scale pretraining with transformer architectures becomes popular, with BERT [12] scaling up to 300 million parameters, followed by GPT-2 [52] at 1.5 billion parameters.”From Megatron-Turing NLG · §Related Works
- BLIP2022 · cited 2×, 1 in Method“The text encoder is the same as BERT (Devlin et al. 2019), where a [CLS] token is appended to the beginning of the text input to summarize the sentence.”From BLIP · §Method
- PaLM2022 · cited 3דBroadly, language modeling refers to approaches for predicting either the next token in a sequence or for predicting masked spans (Devlin et al. 2019; Raffel et al. 2020).”From PaLM · §Related Work
- Flamingo2022 · cited 2דThe paradigm of first pretraining on a vast amount of data followed by an adaptation on a downstream task has become standard [75, 32, 52, 44, 23, 87, 108, 11].”From Flamingo · §Related work
- CoCa2022 · cited 2דPretraining ConvNets [18] or Transformers [19] on large-scale annotated data such as ImageNet [6, 7, 8], Instagram [20] or JFT [21] has become a popular strategy towards solving visual recognition problems including clas…”From CoCa · §Related Work
- VL-BEiT2022 · cited 3×, 1 in Method“Following BERT [7], we randomly mask 15% tokens of monomodal text data.”From VL-BEiT · §Methods
- BIG-bench2022 · cited 4דTask cites: (Devlin et al. 2018; Wu et al. 2016; Won et al. 2021; Xu et al. 2018; Edmiston & Stratos 2018; El-Kishky et al. 2019)”From BIG-bench · §Author contributions
- BEiT v22022 · cited 2דThe MIM method has achieved great success in language task (Devlin et al. 2019; Dong et al. 2019; Bao et al. 2020).”From BEiT v2 · §Related Work
- BEiT-32022 · cited 3דSecond, the pretraining task based on masked data modeling has been successfully applied to various modalities, such as texts [14], images [3, 40], and image-text pairs [7].”From BEiT-3 · §Introduction: The Big Convergence
- U-PaLM2022 · cited 2דWhile there have been many paradigms and self-supervision methods proposed to train these models (Devlin et al. 2018; Clark et al. 2020b; Yang et al. 2019; Raffel et al. 2019), to this date most large language models (i.…”From U-PaLM · §Related Work
- BLIP-22023 · cited 1×, 1 in Method“We initialize Q-Former with the pre-trained weights of BERTbase (Devlin et al. 2019), whereas the cross-attention layers are randomly initialized.”From BLIP-2 · §Method
Abstract
We introduce a new language representation model called BERT, which stands for Bidirectional Encoder Representations from Transformers. Unlike recent language representation models, BERT is designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers. As a result, the pre-trained BERT model can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks, such as question answering and language inference, without substantial task-specific architecture modifications. BERT is conceptually simple and empirically powerful. It obtains new state-of-the-art results on eleven natural language processing tasks, including pushing the GLUE score to 80.5% (7.7% point absolute improvement), MultiNLI accuracy to 86.7% (4.6% absolute improvement), SQuAD v1.1 question answering Test F1 to 93.2 (1.5 point absolute improvement) and SQuAD v2.0 Test F1 to 83.1 (5.1 point absolute improvement).