Paper Lineage
Esc
MethodAug 2019arXiv 1908.03557cs.CV

VisualBERT: A Simple and Performant Baseline for Vision and Language

Liunian Harold Li, Mark Yatskar, Da Yin and 2 others

We propose VisualBERT, a simple and flexible framework for modeling a broad range of vision-and-language tasks. VisualBERT consists of a stack of Transformer layers that implicitly align elements of an input text and regions in an associated input image with self-attention.

From the abstract

Built on

9 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (9)

  • BERT2018 · cited 5×, 1 in Method
    “Our work is inspired by BERT (Devlin et al. 2019), a Transformer-based representation model for natural language.”
    From this paper · §Related Work
  • COCO Captions2015 · cited 3×, 1 in Method
    “Therefore we reach to a resource of paired data: COCO (Chen et al. 2015) that contains images each paired with 5 independent captions.”
    From this paper · §A Joint Representation Model for Vision and Language
  • Transformer2017 · cited 2×, 1 in Method
    “BERT (Devlin et al. 2019) is a Transformer (Vaswani et al. 2017) with subwords (Wu et al. 2016) as input and trained using language modeling objectives.”
    From this paper · §A Joint Representation Model for Vision and Language
  • GNMT2016 · cited 1×, 1 in Method
    “BERT (Devlin et al. 2019) is a Transformer (Vaswani et al. 2017) with subwords (Wu et al. 2016) as input and trained using language modeling objectives.”
    From this paper · §A Joint Representation Model for Vision and Language
Show 5 more
  • VQA v22016 · cited 4×
    “We evaluate VisualBERT on four different types of vision-and-language applications: (1) Visual Question Answering (VQA 2.0) (Goyal et al. 2017), (2) Visual Commonsense Reasoning (VCR) (Zellers et al. 2019), (3) Natural L…”
    From this paper · §Experiment
  • NLVR22018 · cited 4×
    “We evaluate VisualBERT on four different types of vision-and-language applications: (1) Visual Question Answering (VQA 2.0) (Goyal et al. 2017), (2) Visual Commonsense Reasoning (VCR) (Zellers et al. 2019), (3) Natural L…”
    From this paper · §Experiment
  • Flickr30k Entities2015 · cited 3×
    “Various tasks such as visual question answering (Antol et al. 2015; Goyal et al. 2017), textual grounding (Kazemzadeh et al. 2014; Plummer et al. 2015), and visual reasoning (Suhr et al. 2019; Zellers et al. 2019) have b…”
    From this paper · §Related Work
  • VQA2015 · cited 2×
    “Various tasks such as visual question answering (Antol et al. 2015; Goyal et al. 2017), textual grounding (Kazemzadeh et al. 2014; Plummer et al. 2015), and visual reasoning (Suhr et al. 2019; Zellers et al. 2019) have b…”
    From this paper · §Related Work
  • “Various tasks such as visual question answering (Antol et al. 2015; Goyal et al. 2017), textual grounding (Kazemzadeh et al. 2014; Plummer et al. 2015), and visual reasoning (Suhr et al. 2019; Zellers et al. 2019) have b…”
    From this paper · §Related Work

Led to

  • UNITER2019 · cited 3×
    “On the other hand, B2T2 [1], VisualBERT [25], Unicoder-VL [24] and VL-BERT [42] proposed the single-stream architecture, where a single Transformer is applied to both images and text.”
    From UNITER · §Related Work
  • 12-in-12019 · cited 4×, 1 in Method
    “The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”
    From 12-in-1 · §Introduction
  • VirTex2020 · cited 2×
    “Inspired by the success of BERT [64] in NLP, several recent methods use Transformers [29] to learn transferable joint representations of images and text [65, 66, 67, 68, 69, 70, 71, 72].”
    From VirTex · §Related Work
  • VL-BERT meta-analysis2020 · cited 2×
    “While most approaches present very similar ways to embed spatial locations, VL-BERT relies on a more complex geometry embedding and they are, instead, missing in VisualBERT Li et al. 2019.”
    From VL-BERT meta-analysis · §Vision-and-Language BERTs
  • ViLT2021 · cited 3×, 2 in Method
    “Backbone: ResNet-101 (Lu et al. 2019; Tan & Bansal 2019; Su et al. 2019) and ResNext-152 (Li et al. 2019; Li et al. 2020a; Zhang et al. 2021) are two commonly used backbones.”
    From ViLT · §Background
  • Conceptual 12M2021 · cited 3×
    “Based directly upon BERT, V+L pre-training research has largely been focused on V+L understanding [55, 49, 21, 77, 3, 74, 48, 56], with classification or regression tasks that do not involve generation.”
    From Conceptual 12M · §Related Work
  • SimVLM2021 · cited 2×
    “A line of work (Tan & Bansal 2019; Lu et al. 2019; Li et al. 2019; Chen et al. 2020b; Li et al. 2020; Su et al. 2020; Zhang et al. 2021) has explored vision-language pretraining (VLP) that learns a joint representation o…”
    From SimVLM · §Introduction
  • METER2021 · cited 6×, 4 in Method
    “Vision-and-language pre-training (VLP) has now become the de facto practice to tackle these tasks tan-bansal-2019-lxmert; li2019visualbert; lu2019vilbert; su2019vl; chen2020uniter; li2020oscar.”
    From METER · §Introduction
Abstract

We propose VisualBERT, a simple and flexible framework for modeling a broad range of vision-and-language tasks. VisualBERT consists of a stack of Transformer layers that implicitly align elements of an input text and regions in an associated input image with self-attention. We further propose two visually-grounded language model objectives for pre-training VisualBERT on image caption data. Experiments on four vision-and-language tasks including VQA, VCR, NLVR2, and Flickr30K show that VisualBERT outperforms or rivals with state-of-the-art models while being significantly simpler. Further analysis demonstrates that VisualBERT can ground elements of language to image regions without any explicit supervision and is even sensitive to syntactic relationships, tracking, for example, associations between verbs and image regions corresponding to their arguments.