ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision
Vision-and-Language Pre-training (VLP) has improved performance on various joint vision-and-language downstream tasks. , ResNet).
Also cited · not yet reviewed (20)
- LXMERT2019 · cited 4×, 2 in Method“Backbone: ResNet-101 (Lu et al. 2019; Tan & Bansal 2019; Su et al. 2019) and ResNext-152 (Li et al. 2019; Li et al. 2020a; Zhang et al. 2021) are two commonly used backbones.”From this paper · §Background
- VinVL2021 · cited 4×, 2 in Method“Following OSCAR (Li et al. 2020b) and VinVL (Zhang et al. 2021), we use the pair method.”From this paper · §Experiments
- Bottom-Up Top-Down attention2017 · cited 3×, 2 in Method“Most VLP models employ an object detector pre-trained on the Visual Genome dataset (Krishna et al. 2017) annotated with 1,600 object classes and 400 attribute classes as in Anderson et al. 2018.”From this paper · §Introduction
- NLVR22018 · cited 3×, 2 in Method“We evaluate ViLT on two widely explored types of vision-and-language downstream tasks: for classification, we use VQAv2 (Goyal et al. 2017) and NLVR2 (Suhr et al. 2018), and for retrieval, we use MSCOCO and Flickr30K (F3…”From this paper · §Experiments
Show 16 more
- ViLBERT2019 · cited 3×, 2 in Method“Backbone: ResNet-101 (Lu et al. 2019; Tan & Bansal 2019; Su et al. 2019) and ResNext-152 (Li et al. 2019; Li et al. 2020a; Zhang et al. 2021) are two commonly used backbones.”From this paper · §Background
- VisualBERT2019 · cited 3×, 2 in Method“Backbone: ResNet-101 (Lu et al. 2019; Tan & Bansal 2019; Su et al. 2019) and ResNext-152 (Li et al. 2019; Li et al. 2020a; Zhang et al. 2021) are two commonly used backbones.”From this paper · §Background
- Grid features for VQA2020 · cited 2×, 2 in Method“Applying NMS to each and every class becomes a major runtime bottleneck with a large number of classes, e.g. 1.6K in the VG dataset (Jiang et al. 2020).”From this paper · §Background
- VL-BERT meta-analysis2020 · cited 2×, 2 in Method“Bugliarello et al. 2020 classifies interaction schema into two categories: (1) single-stream approaches (e.g., VisualBERT (Li et al. 2019), UNITER (Chen et al. 2019)) where layers collectively operate on a concatenation…”From this paper · §Background
- UNITER2019 · cited 5×, 1 in Method“Plus, inspired by the word region alignment objective in Chen et al. 2019, we design word patch alignment (WPA) that computes the alignment score between two subsets of zDz^{D}: zD|tz^{D}|_{t} (textual subset) and zD|vz^…”From this paper · §Vision-and-Language Transformer
- ViT2020 · cited 4×, 1 in Method“Recent work (Dosovitskiy et al. 2020; Touvron et al. 2020) demonstrated that using a simple linear projection of a patch is effective enough to embed pixels before feeding them into transformers.”From this paper · §Introduction
- Unicoder-VL2019 · cited 2×, 1 in Method“Backbone: ResNet-101 (Lu et al. 2019; Tan & Bansal 2019; Su et al. 2019) and ResNext-152 (Li et al. 2019; Li et al. 2020a; Zhang et al. 2021) are two commonly used backbones.”From this paper · §Background
- VL-BERT2019 · cited 2×, 1 in Method“Backbone: ResNet-101 (Lu et al. 2019; Tan & Bansal 2019; Su et al. 2019) and ResNext-152 (Li et al. 2019; Li et al. 2020a; Zhang et al. 2021) are two commonly used backbones.”From this paper · §Background
- Faster R-CNN2015 · cited 1×, 1 in Method“They are obtained from an off-the-shelf object detector like Faster R-CNN (Ren et al. 2016).”From this paper · §Background
- CLIP2021 · cited 1×, 1 in Method“CLIP (Radford et al. 2021) belongs to Figure 2(b) as it uses separate but equally expensive transformer embedders for each modality.”From this paper · §Background
- BERT2018 · cited 4דFollowing the heuristics of Devlin et al. 2019, we randomly mask tt with the probability of 0.15.”From this paper · §Vision-and-Language Transformer
- Karpathy visual-semantic alignment2014 · cited 2דWe evaluate ViLT on two widely explored types of vision-and-language downstream tasks: for classification, we use VQAv2 (Goyal et al. 2017) and NLVR2 (Suhr et al. 2018), and for retrieval, we use MSCOCO and Flickr30K (F3…”From this paper · §Experiments
- Visual Genome2016 · cited 2דMost VLP models employ an object detector pre-trained on the Visual Genome dataset (Krishna et al. 2017) annotated with 1,600 object classes and 400 attribute classes as in Anderson et al. 2018.”From this paper · §Introduction
- SimCLR2020 · cited 2דPrevious work on contrastive visual representation learning (Chen et al. 2020a; Chen et al. 2020b) showed that gaussian blur, not employed by RandAugment, brings noticeable gains to downstream performance compared with a…”From this paper · §Conclusion and Future Work
- Oscar2020 · cited 2דFollowing OSCAR (Li et al. 2020b) and VinVL (Zhang et al. 2021), we use the pair method.”From this paper · §Experiments
- DeiT2020 · cited 2דRecent work (Dosovitskiy et al. 2020; Touvron et al. 2020) demonstrated that using a simple linear projection of a patch is effective enough to embed pixels before feeding them into transformers.”From this paper · §Introduction
Led to
- ALBEF2021 · cited 2דA recent method [21] improves inference speed by removing the object detector, but results in lower performance.”From ALBEF · §Related Work
- SimVLM2021 · cited 2דSome recent efforts have also explored VLP without object detection module (Xu et al. 2021; Kim et al. 2021; Huang et al. 2021), but they only use clean pretraining data with small scales and thus their zero-shot capabil…”From SimVLM · §Related Work
- VLMo2021 · cited 10דPre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”From VLMo · §Related Work
- METER2021 · cited 7×, 4 in Method“This can lead to several problems: first, the object detectors are not perfect, but are usually kept frozen during VLP, which limits the capacity of the VLP models; second, it is time-consuming to extract region features…”From METER · §Introduction
- Florence2021 · cited 1×, 1 in Method“Recently, there is an increasing trend (Huang et al. 2021; Xue et al. 2021; Wang et al. 2021; Kim et al. 2021; Dou et al. 2021) of end-to-end approaches to reduce dependency on the object bounding box, which instead cons…”From Florence · §Approach
- ViTCAP2021 · cited 6דThese intermediate operations unavoidably cause training inefficiency and high inference latency at prediction stage kim2021vilt; wang2020minivlm; 2) require box annotations and largely limit the flexibility in training…”From ViTCAP · §Introduction
- BLIP2022 · cited 1×, 1 in Method“Compared to using pre-trained object detectors for visual feature extraction (Chen et al. 2020), using a ViT is more computation-friendly and has been adopted by the more recent methods (Li et al. 2021a; Kim et al. 2021)…”From BLIP · §Method
- VL-BEiT2022 · cited 6דFollowing previous work [16, 41], we use VQA 2.0 dataset [10], and formulate the task as a classification problem to choose the answer from 3,1293,129 most frequent answers.”From VL-BEiT · §Experiments
- BEiT-32022 · cited 6×, 1 in Method“In contrast, previous vision-language models [36, 65, 26, 34, 53, 30, 62] usually employ multiple pretraining tasks, such as image-text contrast, image-text matching, and word-patch/region alignment.”From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
Abstract
Vision-and-Language Pre-training (VLP) has improved performance on various joint vision-and-language downstream tasks. Current approaches to VLP heavily rely on image feature extraction processes, most of which involve region supervision (e.g., object detection) and the convolutional architecture (e.g., ResNet). Although disregarded in the literature, we find it problematic in terms of both (1) efficiency/speed, that simply extracting input features requires much more computation than the multimodal interaction steps; and (2) expressive power, as it is upper bounded to the expressive power of the visual embedder and its predefined visual vocabulary. In this paper, we present a minimal VLP model, Vision-and-Language Transformer (ViLT), monolithic in the sense that the processing of visual inputs is drastically simplified to just the same convolution-free manner that we process textual inputs. We show that ViLT is up to tens of times faster than previous VLP models, yet with competitive or better downstream task performance. Our code and pre-trained weights are available at https://github.com/dandelin/vilt.