ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
We present ViLBERT (short for Vision-and-Language BERT), a model for learning task-agnostic joint representations of image content and natural language. We extend the popular BERT architecture to a multi-modal two-stream model, pro-cessing both visual and textual inputs in separate streams that interact through co-attentional transformer layers.
Also cited · not yet reviewed (10)
- Transformer2017 · cited 2×, 2 in Method“These tokens are mapped to learned encodings and passed through LL “encoder-style” transformer blocks [27] to produce final representations h0,…,hTh_{0},\dots,h_{T}.”From this paper · §Approach
- BERT2018 · cited 9×, 1 in Method“In analogy to the training tasks in [12], we train our model on Conceptual Captions on two proxy tasks: predicting the semantics of masked words and image regions given the unmasked inputs, and predicting whether an imag…”From this paper · §Introduction
- Bottom-Up Top-Down attention2017 · cited 3×, 1 in Method“We use Faster R-CNN [31] (with ResNet-101 [11] backbone) pretrained on the Visual Genome dataset [16] (see [30] for details) to extract region features.”From this paper · §Experimental Settings
- GNMT2016 · cited 1×, 1 in Method“For a given token, the input representation is a sum of a token-specific learned embedding [28] and encodings for position (i.e. token’s index in the sequence) and segment (i.e. index of the token’s sentence if multiple…”From this paper · §Approach
Show 6 more
- COCO Captions2015 · cited 3דWhile we address many vision-and-language tasks in Sec. 3.2, we do miss some families of tasks including visually grounded dialog [4, 45], embodied tasks like question answering [7] and instruction following [8], and tex…”From this paper · §Related Work
- VQA2015 · cited 3דWe train and evaluate on the VQA 2.0 dataset [3] consisting of 1.1 million questions about COCO images [5] each with 10 answers.”From this paper · §Experimental Settings
- ELMo2018 · cited 3דSelf-supervised language models on the other hand have resulted in significant improvements over prior work [12, 14, 13, 44].”From this paper · §Related Work
- BookCorpus (books & movies)2015 · cited 2דThis pretrain-then-transfer learning approach to vision-and-language tasks follows naturally from its widespread use in both computer vision and natural language processing where it has become the de facto standard due t…”From this paper · §Introduction
- ResNet2015 · cited 2דWe use Faster R-CNN [31] (with ResNet-101 [11] backbone) pretrained on the Visual Genome dataset [16] (see [30] for details) to extract region features.”From this paper · §Experimental Settings
- Visual Genome2016 · cited 2דThis pretrain-then-transfer learning approach to vision-and-language tasks follows naturally from its widespread use in both computer vision and natural language processing where it has become the de facto standard due t…”From this paper · §Introduction
Led to
- Unicoder-VL2019 · cited 4דHowever, for image RoI based methods like SCAN[\citeauthoryearLee et al.2018], Unicoder-VL and ViLBERT [\citeauthoryearLu et al.2019], the backbone of Faster-RCNN is still not fine-tuned with the whole model during cross…”From Unicoder-VL · §Results and Analysis
- VL-BERT2019 · cited 3דIn ViLBERT (Lu et al. 2019) and LXMERT (Tan & Bansal 2019), which are under review or just got accepted, the network architectures are of two single-modal networks applied on input sentences and images respectively, foll…”From VL-BERT · §Related Work
- Unified VLP2019 · cited 4×, 1 in Method“Inspired by the recent success of pre-trained language models such as BERT [\citeauthoryearDevlin et al.2018] and GPT [\citeauthoryearRadford et al.2018, \citeauthoryearRadford et al.2019], there is a growing interest in…”From Unified VLP · §Introduction
- 12-in-12019 · cited 9×, 4 in Method“The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”From 12-in-1 · §Introduction
- Grid features for VQA2020 · cited 2דAs these separately trained features may not be optimal for joint vision and language understanding, a recent hot topic is to develop jointly pre-trained models li2019visualbert; lu2019vilbert; tan2019lxmert; su2019vl; z…”From Grid features for VQA · §Related Work
- VILLA2020 · cited 3דInspired by the success of BERT [13] on natural language understanding, there has been a surging research interest in developing multimodal pre-training methods for vision-and-language representation learning (e.g., ViLB…”From VILLA · §Introduction
- VL-BERT meta-analysis2020 · cited 5דViLBERT (Lu et al. 2019), LXMERT (Tan and Bansal 2019), and ERNIE-ViL (Yu et al. 2021)33 3 ERNIE-ViL uses the dual-stream ViLBERT encoder. are based on a dual-stream paradigm.”From VL-BERT meta-analysis · §Vision-and-Language BERTs
- VL-T52021 · cited 7×, 3 in Method“As shown in Fig.3 (a), existing methods (Tan & Bansal 2019; Lu et al. 2019; Chen et al. 2020) typically introduce a multi-layer perceptron (MLP) multi-label classifier head on top of h[CLS]xh^{x}_{\texttt{[CLS]}}, which…”From VL-T5 · §Model
- ViLT2021 · cited 3×, 2 in Method“Backbone: ResNet-101 (Lu et al. 2019; Tan & Bansal 2019; Su et al. 2019) and ResNext-152 (Li et al. 2019; Li et al. 2020a; Zhang et al. 2021) are two commonly used backbones.”From ViLT · §Background
- ALIGN2021 · cited 2דRecently more advanced models emerge with cross-modal attention layers (Liu et al. 2019a; Lu et al. 2019; Chen et al. 2020c; Huang et al. 2020b) and show superior performance in image-text matching tasks.”From ALIGN · §Related Work
- Conceptual 12M2021 · cited 12×, 5 in Method“The large scale nature and the high degree of textual and visual diversity make this dataset particularly suited to V+L pre-training [55, 21, 74, 88, 48, 56].”From Conceptual 12M · §Vision-and-Language Pre-Training Data
- SimVLM2021 · cited 2דOn the other hand, multiple cross-modality loss functions have been proposed as part of the training objectives, for example image-text matching (Tan & Bansal 2019; Lu et al. 2019; Xu et al. 2021), masked region classifi…”From SimVLM · §Related Work
- VLMo2021 · cited 4דThis leads to a significant speedup compared with previous models using image region features, which are extracted by an off-the-shelf object detector [30, 41, 3, 14, 25, 49].”From VLMo · §Experiments
- METER2021 · cited 6×, 4 in Method“Vision-and-language pre-training (VLP) has now become the de facto practice to tackle these tasks tan-bansal-2019-lxmert; li2019visualbert; lu2019vilbert; su2019vl; chen2020uniter; li2020oscar.”From METER · §Introduction
- Flamingo2022 · cited 2דIn particular, BERT [23] inspired a large body of vision-language work [66, 106, 16, 38, 121, 61, 109, 151, 118, 59, 29, 28, 142, 143, 101, 107].”From Flamingo · §Related work
- VL-BEiT2022 · cited 2דPrevious models [24, 35, 21, 46] use an off-the-shelf object detector like Faster R-CNN [30] to obtain image region features.”From VL-BEiT · §Related Work
Abstract
We present ViLBERT (short for Vision-and-Language BERT), a model for learning task-agnostic joint representations of image content and natural language. We extend the popular BERT architecture to a multi-modal two-stream model, pro-cessing both visual and textual inputs in separate streams that interact through co-attentional transformer layers. We pretrain our model through two proxy tasks on the large, automatically collected Conceptual Captions dataset and then transfer it to multiple established vision-and-language tasks -- visual question answering, visual commonsense reasoning, referring expressions, and caption-based image retrieval -- by making only minor additions to the base architecture. We observe significant improvements across tasks compared to existing task-specific models -- achieving state-of-the-art on all four tasks. Our work represents a shift away from learning groundings between vision and language only as part of task training and towards treating visual grounding as a pretrainable and transferable capability.