Paper Lineage
Esc
MethodAug 2019arXiv 1908.08530cs.CV

VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Weijie Su, Xizhou Zhu, Yue Cao and 4 others

We introduce a new pre-trainable generic representation for visual-linguistic tasks, called Visual-Linguistic BERT (VL-BERT for short). VL-BERT adopts the simple yet powerful Transformer model as the backbone, and extends it to take both visual and linguistic embedded features as input.

From the abstract

Built on

11 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (11)

  • BERT2018 · cited 6×
    “After that, a serious of approaches are proposed for pre-training the generic representation, mainly based on Transformers, such as GPT (Radford et al. 2018), BERT (Devlin et al. 2018), GPT-2 (Radford et al. 2019), XLNet…”
    From this paper · §Related Work
  • Transformer2017 · cited 5×
    “In the milestone work of Transformers (Vaswani et al. 2017), the Transformer attention module is proposed as a generic building block for various NLP tasks.”
    From this paper · §Related Work
  • “To better exploit the generic representation, we pre-train VL-BERT at both large visual-linguistic corpus and text-only datasets11 1 Here we exploit the Conceptual Captions dataset (Sharma et al. 2018) as the visual-ling…”
    From this paper · §Introduction
  • VQA v22016 · cited 3×
    “Here we conduct experiments on the widely-used VQA v2.0 dataset (Goyal et al. 2017), which is built based on the COCO (Lin et al. 2014) images.”
    From this paper · §Experiment
Show 7 more
  • “Visual content embedding is produced by Faster R-CNN + ResNet-101, initialized from parameters pre-trained on Visual Genome (Krishna et al. 2017) for object detection (see BUTD (Anderson et al. 2018)).”
    From this paper · §Experiment
  • ViLBERT2019 · cited 3×
    “In ViLBERT (Lu et al. 2019) and LXMERT (Tan & Bansal 2019), which are under review or just got accepted, the network architectures are of two single-modal networks applied on input sentences and images respectively, foll…”
    From this paper · §Related Work
  • LXMERT2019 · cited 3×
    “In ViLBERT (Lu et al. 2019) and LXMERT (Tan & Bansal 2019), which are under review or just got accepted, the network architectures are of two single-modal networks applied on input sentences and images respectively, foll…”
    From this paper · §Related Work
  • MS COCO2014 · cited 2×
    “Here we conduct experiments on the widely-used VQA v2.0 dataset (Goyal et al. 2017), which is built based on the COCO (Lin et al. 2014) images.”
    From this paper · §Experiment
  • VQA2015 · cited 2×
    “Meanwhile, for tasks at the intersection of vision and language, such as image captioning (Young et al. 2014; Chen et al. 2015; Sharma et al. 2018), visual question answering (VQA) (Antol et al. 2015; Johnson et al. 2017…”
    From this paper · §Introduction
  • Visual Genome2016 · cited 2×
    “Visual content embedding is produced by Faster R-CNN + ResNet-101, initialized from parameters pre-trained on Visual Genome (Krishna et al. 2017) for object detection (see BUTD (Anderson et al. 2018)).”
    From this paper · §Experiment
  • GNMT2016 · cited 2×
    “Token Embedding Following the practice in BERT, the linguistic words are embedded with WordPiece embeddings (Wu et al. 2016) with a 30,000 vocabulary.”
    From this paper · §VL-BERT

Led to

  • 12-in-12019 · cited 5×, 1 in Method
    “The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”
    From 12-in-1 · §Introduction
  • Oscar2020 · cited 2×
    “The existing methods [37, 38, 22, 5, 46, 35, 19, 10] employ BERT-like objectives [6] to learn cross-modal representations from a concatenated-sequence of visual region features and language token embeddings.”
    From Oscar · §Related Work
  • VILLA2020 · cited 2×
    “Compared to this two-stream architecture, recent work such as VL-BERT [60], VisualBERT [33], B2T2 [1], Unicoder-VL [30] and UNITER [12] advocate a single-stream model design, where two modalities are directly fused in ea…”
    From VILLA · §Related Work
  • ConVIRT2020 · cited 2×
    “Contrastive-Binary-Loss: This baseline differs from ConVIRT by contrasting the paired image and text representations with a binary classification head, as is widely done in visual-linguistic pretraining work Tan and Bans…”
    From ConVIRT · §Experiments
  • VL-BERT meta-analysis2020 · cited 3×
    “While most V&L BERTs follow this paradigm, some studies find beneficial to jointly learn the visual encoder with language Su et al. 2020; Huang et al. 2020; Radford et al. 2021; Kim et al. 2021.”
    From VL-BERT meta-analysis · §Results
  • ViLT2021 · cited 2×, 1 in Method
    “Backbone: ResNet-101 (Lu et al. 2019; Tan & Bansal 2019; Su et al. 2019) and ResNext-152 (Li et al. 2019; Li et al. 2020a; Zhang et al. 2021) are two commonly used backbones.”
    From ViLT · §Background
  • Conceptual 12M2021 · cited 6×, 1 in Method
    “The large scale nature and the high degree of textual and visual diversity make this dataset particularly suited to V+L pre-training [55, 21, 74, 88, 48, 56].”
    From Conceptual 12M · §Vision-and-Language Pre-Training Data
  • SimVLM2021 · cited 2×
    “While a variety of approaches have been proposed, a large portion of them require object detection for image region feature regression or tagging as part of the pre-training objectives (Tan & Bansal 2019; Su et al. 2020;…”
    From SimVLM · §Related Work
  • VLMo2021 · cited 5×
    “Pre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”
    From VLMo · §Related Work
  • METER2021 · cited 4×, 2 in Method
    “Vision-and-language pre-training (VLP) has now become the de facto practice to tackle these tasks tan-bansal-2019-lxmert; li2019visualbert; lu2019vilbert; su2019vl; chen2020uniter; li2020oscar.”
    From METER · §Introduction
  • VL-BEiT2022 · cited 3×
    “Previous models [24, 35, 21, 46] use an off-the-shelf object detector like Faster R-CNN [30] to obtain image region features.”
    From VL-BEiT · §Related Work
Abstract

We introduce a new pre-trainable generic representation for visual-linguistic tasks, called Visual-Linguistic BERT (VL-BERT for short). VL-BERT adopts the simple yet powerful Transformer model as the backbone, and extends it to take both visual and linguistic embedded features as input. In it, each element of the input is either of a word from the input sentence, or a region-of-interest (RoI) from the input image. It is designed to fit for most of the visual-linguistic downstream tasks. To better exploit the generic representation, we pre-train VL-BERT on the massive-scale Conceptual Captions dataset, together with text-only corpus. Extensive empirical analysis demonstrates that the pre-training procedure can better align the visual-linguistic clues and benefit the downstream tasks, such as visual commonsense reasoning, visual question answering and referring expression comprehension. It is worth noting that VL-BERT achieved the first place of single model on the leaderboard of the VCR benchmark. Code is released at \url{https://github.com/jackroos/VL-BERT}.