Paper Lineage
Esc
MethodFeb 2021arXiv 2102.03334stat.ML

ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision

Wonjae Kim, Bokyung Son, Ildoo Kim

Vision-and-Language Pre-training (VLP) has improved performance on various joint vision-and-language downstream tasks. , ResNet).

From the abstract

Built on

20 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (20)

  • LXMERT2019 · cited 4×, 2 in Method
    “Backbone: ResNet-101 (Lu et al. 2019; Tan & Bansal 2019; Su et al. 2019) and ResNext-152 (Li et al. 2019; Li et al. 2020a; Zhang et al. 2021) are two commonly used backbones.”
    From this paper · §Background
  • VinVL2021 · cited 4×, 2 in Method
    “Following OSCAR (Li et al. 2020b) and VinVL (Zhang et al. 2021), we use the pair method.”
    From this paper · §Experiments
  • Bottom-Up Top-Down attention2017 · cited 3×, 2 in Method
    “Most VLP models employ an object detector pre-trained on the Visual Genome dataset (Krishna et al. 2017) annotated with 1,600 object classes and 400 attribute classes as in Anderson et al. 2018.”
    From this paper · §Introduction
  • NLVR22018 · cited 3×, 2 in Method
    “We evaluate ViLT on two widely explored types of vision-and-language downstream tasks: for classification, we use VQAv2 (Goyal et al. 2017) and NLVR2 (Suhr et al. 2018), and for retrieval, we use MSCOCO and Flickr30K (F3…”
    From this paper · §Experiments
Show 16 more
  • ViLBERT2019 · cited 3×, 2 in Method
    “Backbone: ResNet-101 (Lu et al. 2019; Tan & Bansal 2019; Su et al. 2019) and ResNext-152 (Li et al. 2019; Li et al. 2020a; Zhang et al. 2021) are two commonly used backbones.”
    From this paper · §Background
  • VisualBERT2019 · cited 3×, 2 in Method
    “Backbone: ResNet-101 (Lu et al. 2019; Tan & Bansal 2019; Su et al. 2019) and ResNext-152 (Li et al. 2019; Li et al. 2020a; Zhang et al. 2021) are two commonly used backbones.”
    From this paper · §Background
  • Grid features for VQA2020 · cited 2×, 2 in Method
    “Applying NMS to each and every class becomes a major runtime bottleneck with a large number of classes, e.g. 1.6K in the VG dataset (Jiang et al. 2020).”
    From this paper · §Background
  • VL-BERT meta-analysis2020 · cited 2×, 2 in Method
    “Bugliarello et al. 2020 classifies interaction schema into two categories: (1) single-stream approaches (e.g., VisualBERT (Li et al. 2019), UNITER (Chen et al. 2019)) where layers collectively operate on a concatenation…”
    From this paper · §Background
  • UNITER2019 · cited 5×, 1 in Method
    “Plus, inspired by the word region alignment objective in Chen et al. 2019, we design word patch alignment (WPA) that computes the alignment score between two subsets of zDz^{D}: zD|tz^{D}|_{t} (textual subset) and zD|vz^…”
    From this paper · §Vision-and-Language Transformer
  • ViT2020 · cited 4×, 1 in Method
    “Recent work (Dosovitskiy et al. 2020; Touvron et al. 2020) demonstrated that using a simple linear projection of a patch is effective enough to embed pixels before feeding them into transformers.”
    From this paper · §Introduction
  • Unicoder-VL2019 · cited 2×, 1 in Method
    “Backbone: ResNet-101 (Lu et al. 2019; Tan & Bansal 2019; Su et al. 2019) and ResNext-152 (Li et al. 2019; Li et al. 2020a; Zhang et al. 2021) are two commonly used backbones.”
    From this paper · §Background
  • VL-BERT2019 · cited 2×, 1 in Method
    “Backbone: ResNet-101 (Lu et al. 2019; Tan & Bansal 2019; Su et al. 2019) and ResNext-152 (Li et al. 2019; Li et al. 2020a; Zhang et al. 2021) are two commonly used backbones.”
    From this paper · §Background
  • Faster R-CNN2015 · cited 1×, 1 in Method
    “They are obtained from an off-the-shelf object detector like Faster R-CNN (Ren et al. 2016).”
    From this paper · §Background
  • CLIP2021 · cited 1×, 1 in Method
    “CLIP (Radford et al. 2021) belongs to Figure 2(b) as it uses separate but equally expensive transformer embedders for each modality.”
    From this paper · §Background
  • BERT2018 · cited 4×
    “Following the heuristics of Devlin et al. 2019, we randomly mask tt with the probability of 0.15.”
    From this paper · §Vision-and-Language Transformer
  • “We evaluate ViLT on two widely explored types of vision-and-language downstream tasks: for classification, we use VQAv2 (Goyal et al. 2017) and NLVR2 (Suhr et al. 2018), and for retrieval, we use MSCOCO and Flickr30K (F3…”
    From this paper · §Experiments
  • Visual Genome2016 · cited 2×
    “Most VLP models employ an object detector pre-trained on the Visual Genome dataset (Krishna et al. 2017) annotated with 1,600 object classes and 400 attribute classes as in Anderson et al. 2018.”
    From this paper · §Introduction
  • SimCLR2020 · cited 2×
    “Previous work on contrastive visual representation learning (Chen et al. 2020a; Chen et al. 2020b) showed that gaussian blur, not employed by RandAugment, brings noticeable gains to downstream performance compared with a…”
    From this paper · §Conclusion and Future Work
  • Oscar2020 · cited 2×
    “Following OSCAR (Li et al. 2020b) and VinVL (Zhang et al. 2021), we use the pair method.”
    From this paper · §Experiments
  • DeiT2020 · cited 2×
    “Recent work (Dosovitskiy et al. 2020; Touvron et al. 2020) demonstrated that using a simple linear projection of a patch is effective enough to embed pixels before feeding them into transformers.”
    From this paper · §Introduction

Led to

  • ALBEF2021 · cited 2×
    “A recent method [21] improves inference speed by removing the object detector, but results in lower performance.”
    From ALBEF · §Related Work
  • SimVLM2021 · cited 2×
    “Some recent efforts have also explored VLP without object detection module (Xu et al. 2021; Kim et al. 2021; Huang et al. 2021), but they only use clean pretraining data with small scales and thus their zero-shot capabil…”
    From SimVLM · §Related Work
  • VLMo2021 · cited 10×
    “Pre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”
    From VLMo · §Related Work
  • METER2021 · cited 7×, 4 in Method
    “This can lead to several problems: first, the object detectors are not perfect, but are usually kept frozen during VLP, which limits the capacity of the VLP models; second, it is time-consuming to extract region features…”
    From METER · §Introduction
  • Florence2021 · cited 1×, 1 in Method
    “Recently, there is an increasing trend (Huang et al. 2021; Xue et al. 2021; Wang et al. 2021; Kim et al. 2021; Dou et al. 2021) of end-to-end approaches to reduce dependency on the object bounding box, which instead cons…”
    From Florence · §Approach
  • ViTCAP2021 · cited 6×
    “These intermediate operations unavoidably cause training inefficiency and high inference latency at prediction stage kim2021vilt; wang2020minivlm; 2) require box annotations and largely limit the flexibility in training…”
    From ViTCAP · §Introduction
  • BLIP2022 · cited 1×, 1 in Method
    “Compared to using pre-trained object detectors for visual feature extraction (Chen et al. 2020), using a ViT is more computation-friendly and has been adopted by the more recent methods (Li et al. 2021a; Kim et al. 2021)…”
    From BLIP · §Method
  • VL-BEiT2022 · cited 6×
    “Following previous work [16, 41], we use VQA 2.0 dataset [10], and formulate the task as a classification problem to choose the answer from 3,1293,129 most frequent answers.”
    From VL-BEiT · §Experiments
  • BEiT-32022 · cited 6×, 1 in Method
    “In contrast, previous vision-language models [36, 65, 26, 34, 53, 30, 62] usually employ multiple pretraining tasks, such as image-text contrast, image-text matching, and word-patch/region alignment.”
    From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
Abstract

Vision-and-Language Pre-training (VLP) has improved performance on various joint vision-and-language downstream tasks. Current approaches to VLP heavily rely on image feature extraction processes, most of which involve region supervision (e.g., object detection) and the convolutional architecture (e.g., ResNet). Although disregarded in the literature, we find it problematic in terms of both (1) efficiency/speed, that simply extracting input features requires much more computation than the multimodal interaction steps; and (2) expressive power, as it is upper bounded to the expressive power of the visual embedder and its predefined visual vocabulary. In this paper, we present a minimal VLP model, Vision-and-Language Transformer (ViLT), monolithic in the sense that the processing of visual inputs is drastically simplified to just the same convolution-free manner that we process textual inputs. We show that ViLT is up to tens of times faster than previous VLP models, yet with competitive or better downstream task performance. Our code and pre-trained weights are available at https://github.com/dandelin/vilt.