Paper Lineage
Esc
MethodAug 2019arXiv 1908.02265cs.CV

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Jiasen Lu, Dhruv Batra, Devi Parikh, Stefan Lee

We present ViLBERT (short for Vision-and-Language BERT), a model for learning task-agnostic joint representations of image content and natural language. We extend the popular BERT architecture to a multi-modal two-stream model, pro-cessing both visual and textual inputs in separate streams that interact through co-attentional transformer layers.

From the abstract

Built on

10 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (10)

  • Transformer2017 · cited 2×, 2 in Method
    “These tokens are mapped to learned encodings and passed through LL “encoder-style” transformer blocks [27] to produce final representations h0,…,hTh_{0},\dots,h_{T}.”
    From this paper · §Approach
  • BERT2018 · cited 9×, 1 in Method
    “In analogy to the training tasks in [12], we train our model on Conceptual Captions on two proxy tasks: predicting the semantics of masked words and image regions given the unmasked inputs, and predicting whether an imag…”
    From this paper · §Introduction
  • Bottom-Up Top-Down attention2017 · cited 3×, 1 in Method
    “We use Faster R-CNN [31] (with ResNet-101 [11] backbone) pretrained on the Visual Genome dataset [16] (see [30] for details) to extract region features.”
    From this paper · §Experimental Settings
  • GNMT2016 · cited 1×, 1 in Method
    “For a given token, the input representation is a sum of a token-specific learned embedding [28] and encodings for position (i.e. token’s index in the sequence) and segment (i.e. index of the token’s sentence if multiple…”
    From this paper · §Approach
Show 6 more
  • COCO Captions2015 · cited 3×
    “While we address many vision-and-language tasks in Sec. 3.2, we do miss some families of tasks including visually grounded dialog [4, 45], embodied tasks like question answering [7] and instruction following [8], and tex…”
    From this paper · §Related Work
  • VQA2015 · cited 3×
    “We train and evaluate on the VQA 2.0 dataset [3] consisting of 1.1 million questions about COCO images [5] each with 10 answers.”
    From this paper · §Experimental Settings
  • ELMo2018 · cited 3×
    “Self-supervised language models on the other hand have resulted in significant improvements over prior work [12, 14, 13, 44].”
    From this paper · §Related Work
  • “This pretrain-then-transfer learning approach to vision-and-language tasks follows naturally from its widespread use in both computer vision and natural language processing where it has become the de facto standard due t…”
    From this paper · §Introduction
  • ResNet2015 · cited 2×
    “We use Faster R-CNN [31] (with ResNet-101 [11] backbone) pretrained on the Visual Genome dataset [16] (see [30] for details) to extract region features.”
    From this paper · §Experimental Settings
  • Visual Genome2016 · cited 2×
    “This pretrain-then-transfer learning approach to vision-and-language tasks follows naturally from its widespread use in both computer vision and natural language processing where it has become the de facto standard due t…”
    From this paper · §Introduction

Led to

  • Unicoder-VL2019 · cited 4×
    “However, for image RoI based methods like SCAN[\citeauthoryearLee et al.2018], Unicoder-VL and ViLBERT [\citeauthoryearLu et al.2019], the backbone of Faster-RCNN is still not fine-tuned with the whole model during cross…”
    From Unicoder-VL · §Results and Analysis
  • VL-BERT2019 · cited 3×
    “In ViLBERT (Lu et al. 2019) and LXMERT (Tan & Bansal 2019), which are under review or just got accepted, the network architectures are of two single-modal networks applied on input sentences and images respectively, foll…”
    From VL-BERT · §Related Work
  • Unified VLP2019 · cited 4×, 1 in Method
    “Inspired by the recent success of pre-trained language models such as BERT [\citeauthoryearDevlin et al.2018] and GPT [\citeauthoryearRadford et al.2018, \citeauthoryearRadford et al.2019], there is a growing interest in…”
    From Unified VLP · §Introduction
  • 12-in-12019 · cited 9×, 4 in Method
    “The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”
    From 12-in-1 · §Introduction
  • Grid features for VQA2020 · cited 2×
    “As these separately trained features may not be optimal for joint vision and language understanding, a recent hot topic is to develop jointly pre-trained models li2019visualbert; lu2019vilbert; tan2019lxmert; su2019vl; z…”
    From Grid features for VQA · §Related Work
  • VILLA2020 · cited 3×
    “Inspired by the success of BERT [13] on natural language understanding, there has been a surging research interest in developing multimodal pre-training methods for vision-and-language representation learning (e.g., ViLB…”
    From VILLA · §Introduction
  • VL-BERT meta-analysis2020 · cited 5×
    “ViLBERT (Lu et al. 2019), LXMERT (Tan and Bansal 2019), and ERNIE-ViL (Yu et al. 2021)33 3 ERNIE-ViL uses the dual-stream ViLBERT encoder. are based on a dual-stream paradigm.”
    From VL-BERT meta-analysis · §Vision-and-Language BERTs
  • VL-T52021 · cited 7×, 3 in Method
    “As shown in Fig.3 (a), existing methods (Tan & Bansal 2019; Lu et al. 2019; Chen et al. 2020) typically introduce a multi-layer perceptron (MLP) multi-label classifier head on top of h[CLS]xh^{x}_{\texttt{[CLS]}}, which…”
    From VL-T5 · §Model
  • ViLT2021 · cited 3×, 2 in Method
    “Backbone: ResNet-101 (Lu et al. 2019; Tan & Bansal 2019; Su et al. 2019) and ResNext-152 (Li et al. 2019; Li et al. 2020a; Zhang et al. 2021) are two commonly used backbones.”
    From ViLT · §Background
  • ALIGN2021 · cited 2×
    “Recently more advanced models emerge with cross-modal attention layers (Liu et al. 2019a; Lu et al. 2019; Chen et al. 2020c; Huang et al. 2020b) and show superior performance in image-text matching tasks.”
    From ALIGN · §Related Work
  • Conceptual 12M2021 · cited 12×, 5 in Method
    “The large scale nature and the high degree of textual and visual diversity make this dataset particularly suited to V+L pre-training [55, 21, 74, 88, 48, 56].”
    From Conceptual 12M · §Vision-and-Language Pre-Training Data
  • SimVLM2021 · cited 2×
    “On the other hand, multiple cross-modality loss functions have been proposed as part of the training objectives, for example image-text matching (Tan & Bansal 2019; Lu et al. 2019; Xu et al. 2021), masked region classifi…”
    From SimVLM · §Related Work
  • VLMo2021 · cited 4×
    “This leads to a significant speedup compared with previous models using image region features, which are extracted by an off-the-shelf object detector [30, 41, 3, 14, 25, 49].”
    From VLMo · §Experiments
  • METER2021 · cited 6×, 4 in Method
    “Vision-and-language pre-training (VLP) has now become the de facto practice to tackle these tasks tan-bansal-2019-lxmert; li2019visualbert; lu2019vilbert; su2019vl; chen2020uniter; li2020oscar.”
    From METER · §Introduction
  • Flamingo2022 · cited 2×
    “In particular, BERT [23] inspired a large body of vision-language work [66, 106, 16, 38, 121, 61, 109, 151, 118, 59, 29, 28, 142, 143, 101, 107].”
    From Flamingo · §Related work
  • VL-BEiT2022 · cited 2×
    “Previous models [24, 35, 21, 46] use an off-the-shelf object detector like Faster R-CNN [30] to obtain image region features.”
    From VL-BEiT · §Related Work
Abstract

We present ViLBERT (short for Vision-and-Language BERT), a model for learning task-agnostic joint representations of image content and natural language. We extend the popular BERT architecture to a multi-modal two-stream model, pro-cessing both visual and textual inputs in separate streams that interact through co-attentional transformer layers. We pretrain our model through two proxy tasks on the large, automatically collected Conceptual Captions dataset and then transfer it to multiple established vision-and-language tasks -- visual question answering, visual commonsense reasoning, referring expressions, and caption-based image retrieval -- by making only minor additions to the base architecture. We observe significant improvements across tasks compared to existing task-specific models -- achieving state-of-the-art on all four tasks. Our work represents a shift away from learning groundings between vision and language only as part of task training and towards treating visual grounding as a pretrainable and transferable capability.