VL-BERT: Pre-training of Generic Visual-Linguistic Representations
We introduce a new pre-trainable generic representation for visual-linguistic tasks, called Visual-Linguistic BERT (VL-BERT for short). VL-BERT adopts the simple yet powerful Transformer model as the backbone, and extends it to take both visual and linguistic embedded features as input.
Also cited · not yet reviewed (11)
- BERT2018 · cited 6דAfter that, a serious of approaches are proposed for pre-training the generic representation, mainly based on Transformers, such as GPT (Radford et al. 2018), BERT (Devlin et al. 2018), GPT-2 (Radford et al. 2019), XLNet…”From this paper · §Related Work
- Transformer2017 · cited 5דIn the milestone work of Transformers (Vaswani et al. 2017), the Transformer attention module is proposed as a generic building block for various NLP tasks.”From this paper · §Related Work
- BookCorpus (books & movies)2015 · cited 3דTo better exploit the generic representation, we pre-train VL-BERT at both large visual-linguistic corpus and text-only datasets11 1 Here we exploit the Conceptual Captions dataset (Sharma et al. 2018) as the visual-ling…”From this paper · §Introduction
- VQA v22016 · cited 3דHere we conduct experiments on the widely-used VQA v2.0 dataset (Goyal et al. 2017), which is built based on the COCO (Lin et al. 2014) images.”From this paper · §Experiment
Show 7 more
- Bottom-Up Top-Down attention2017 · cited 3דVisual content embedding is produced by Faster R-CNN + ResNet-101, initialized from parameters pre-trained on Visual Genome (Krishna et al. 2017) for object detection (see BUTD (Anderson et al. 2018)).”From this paper · §Experiment
- ViLBERT2019 · cited 3דIn ViLBERT (Lu et al. 2019) and LXMERT (Tan & Bansal 2019), which are under review or just got accepted, the network architectures are of two single-modal networks applied on input sentences and images respectively, foll…”From this paper · §Related Work
- LXMERT2019 · cited 3דIn ViLBERT (Lu et al. 2019) and LXMERT (Tan & Bansal 2019), which are under review or just got accepted, the network architectures are of two single-modal networks applied on input sentences and images respectively, foll…”From this paper · §Related Work
- MS COCO2014 · cited 2דHere we conduct experiments on the widely-used VQA v2.0 dataset (Goyal et al. 2017), which is built based on the COCO (Lin et al. 2014) images.”From this paper · §Experiment
- VQA2015 · cited 2דMeanwhile, for tasks at the intersection of vision and language, such as image captioning (Young et al. 2014; Chen et al. 2015; Sharma et al. 2018), visual question answering (VQA) (Antol et al. 2015; Johnson et al. 2017…”From this paper · §Introduction
- Visual Genome2016 · cited 2דVisual content embedding is produced by Faster R-CNN + ResNet-101, initialized from parameters pre-trained on Visual Genome (Krishna et al. 2017) for object detection (see BUTD (Anderson et al. 2018)).”From this paper · §Experiment
- GNMT2016 · cited 2דToken Embedding Following the practice in BERT, the linguistic words are embedded with WordPiece embeddings (Wu et al. 2016) with a 30,000 vocabulary.”From this paper · §VL-BERT
Led to
- 12-in-12019 · cited 5×, 1 in Method“The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”From 12-in-1 · §Introduction
- Oscar2020 · cited 2דThe existing methods [37, 38, 22, 5, 46, 35, 19, 10] employ BERT-like objectives [6] to learn cross-modal representations from a concatenated-sequence of visual region features and language token embeddings.”From Oscar · §Related Work
- VILLA2020 · cited 2דCompared to this two-stream architecture, recent work such as VL-BERT [60], VisualBERT [33], B2T2 [1], Unicoder-VL [30] and UNITER [12] advocate a single-stream model design, where two modalities are directly fused in ea…”From VILLA · §Related Work
- ConVIRT2020 · cited 2דContrastive-Binary-Loss: This baseline differs from ConVIRT by contrasting the paired image and text representations with a binary classification head, as is widely done in visual-linguistic pretraining work Tan and Bans…”From ConVIRT · §Experiments
- VL-BERT meta-analysis2020 · cited 3דWhile most V&L BERTs follow this paradigm, some studies find beneficial to jointly learn the visual encoder with language Su et al. 2020; Huang et al. 2020; Radford et al. 2021; Kim et al. 2021.”From VL-BERT meta-analysis · §Results
- ViLT2021 · cited 2×, 1 in Method“Backbone: ResNet-101 (Lu et al. 2019; Tan & Bansal 2019; Su et al. 2019) and ResNext-152 (Li et al. 2019; Li et al. 2020a; Zhang et al. 2021) are two commonly used backbones.”From ViLT · §Background
- Conceptual 12M2021 · cited 6×, 1 in Method“The large scale nature and the high degree of textual and visual diversity make this dataset particularly suited to V+L pre-training [55, 21, 74, 88, 48, 56].”From Conceptual 12M · §Vision-and-Language Pre-Training Data
- SimVLM2021 · cited 2דWhile a variety of approaches have been proposed, a large portion of them require object detection for image region feature regression or tagging as part of the pre-training objectives (Tan & Bansal 2019; Su et al. 2020;…”From SimVLM · §Related Work
- VLMo2021 · cited 5דPre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”From VLMo · §Related Work
- METER2021 · cited 4×, 2 in Method“Vision-and-language pre-training (VLP) has now become the de facto practice to tackle these tasks tan-bansal-2019-lxmert; li2019visualbert; lu2019vilbert; su2019vl; chen2020uniter; li2020oscar.”From METER · §Introduction
- VL-BEiT2022 · cited 3דPrevious models [24, 35, 21, 46] use an off-the-shelf object detector like Faster R-CNN [30] to obtain image region features.”From VL-BEiT · §Related Work
Abstract
We introduce a new pre-trainable generic representation for visual-linguistic tasks, called Visual-Linguistic BERT (VL-BERT for short). VL-BERT adopts the simple yet powerful Transformer model as the backbone, and extends it to take both visual and linguistic embedded features as input. In it, each element of the input is either of a word from the input sentence, or a region-of-interest (RoI) from the input image. It is designed to fit for most of the visual-linguistic downstream tasks. To better exploit the generic representation, we pre-train VL-BERT on the massive-scale Conceptual Captions dataset, together with text-only corpus. Extensive empirical analysis demonstrates that the pre-training procedure can better align the visual-linguistic clues and benefit the downstream tasks, such as visual commonsense reasoning, visual question answering and referring expression comprehension. It is worth noting that VL-BERT achieved the first place of single model on the leaderboard of the VCR benchmark. Code is released at \url{https://github.com/jackroos/VL-BERT}.