SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
With recent progress in joint modeling of visual and textual representations, Vision-Language Pretraining (VLP) has achieved impressive performance on many multimodal downstream tasks. However, the requirement for expensive annotations including clean image captions and regional labels limits the scalability of existing approaches, and complicates the pretraining procedure with the introduction of multiple dataset-specific objectives.
Also cited · not yet reviewed (19)
- ALIGN2021 · cited 6דMotivated by recent works (Radford et al. 2021; Ramesh et al. 2021; Jia et al. 2021; Tsimpoukelli et al. 2021) that illustrate zero-shot learning in certain image-text tasks, we train our model using large-scale weakly l…”From this paper · §Related Work
- BERT2018 · cited 5דSelf-supervised textual representation learning (Devlin et al. 2018; Radford et al. 2018; Radford et al. 2019; Liu et al. 2019; Yang et al. 2019; Raffel et al. 2019; Brown et al. 2020) based on Transformers (Vaswani et a…”From this paper · §Introduction
- LXMERT2019 · cited 5דWhile a variety of approaches have been proposed, a large portion of them require object detection for image region feature regression or tagging as part of the pre-training objectives (Tan & Bansal 2019; Su et al. 2020;…”From this paper · §Related Work
- UNITER2019 · cited 5דWhile a variety of approaches have been proposed, a large portion of them require object detection for image region feature regression or tagging as part of the pre-training objectives (Tan & Bansal 2019; Su et al. 2020;…”From this paper · §Related Work
Show 15 more
- CoAtNet2021 · cited 5דFollowing Dai et al. 2021, we experiment with using either the first 2/3/4 ResNet Conv blocks, and empirically observe that the 3 conv block setup works best.”From this paper · §Experiments
- T52019 · cited 4דSelf-supervised textual representation learning (Devlin et al. 2018; Radford et al. 2018; Radford et al. 2019; Liu et al. 2019; Yang et al. 2019; Raffel et al. 2019; Brown et al. 2020) based on Transformers (Vaswani et a…”From this paper · §Introduction
- Oscar2020 · cited 4דWhile a variety of approaches have been proposed, a large portion of them require object detection for image region feature regression or tagging as part of the pre-training objectives (Tan & Bansal 2019; Su et al. 2020;…”From this paper · §Related Work
- GPT-32020 · cited 4דSelf-supervised textual representation learning (Devlin et al. 2018; Radford et al. 2018; Radford et al. 2019; Liu et al. 2019; Yang et al. 2019; Raffel et al. 2019; Brown et al. 2020) based on Transformers (Vaswani et a…”From this paper · §Introduction
- ViT2020 · cited 4דWe follow the setup in ViT (Dosovitskiy et al. 2021) to explore 3 variants of SimVLM, namely “Base”, “Large”, and “Huge”, such that each variant follows the same setting as its corresponding ViT variant.”From this paper · §Experiments
- CLIP2021 · cited 4דMotivated by recent works (Radford et al. 2021; Ramesh et al. 2021; Jia et al. 2021; Tsimpoukelli et al. 2021) that illustrate zero-shot learning in certain image-text tasks, we train our model using large-scale weakly l…”From this paper · §Related Work
- VinVL2021 · cited 3דWhile a variety of approaches have been proposed, a large portion of them require object detection for image region feature regression or tagging as part of the pre-training objectives (Tan & Bansal 2019; Su et al. 2020;…”From this paper · §Related Work
- VL-T52021 · cited 3דFirstly, we follow Cho et al. 2021 and evaluate model performance on questions with rare answers in the Karpathy-test split.”From this paper · §Experiments
- VQA v22016 · cited 2דA line of work (Tan & Bansal 2019; Lu et al. 2019; Li et al. 2019; Chen et al. 2020b; Li et al. 2020; Su et al. 2020; Zhang et al. 2021) has explored vision-language pretraining (VLP) that learns a joint representation o…”From this paper · §Introduction
- ViLBERT2019 · cited 2דOn the other hand, multiple cross-modality loss functions have been proposed as part of the training objectives, for example image-text matching (Tan & Bansal 2019; Lu et al. 2019; Xu et al. 2021), masked region classifi…”From this paper · §Related Work
- VisualBERT2019 · cited 2דA line of work (Tan & Bansal 2019; Lu et al. 2019; Li et al. 2019; Chen et al. 2020b; Li et al. 2020; Su et al. 2020; Zhang et al. 2021) has explored vision-language pretraining (VLP) that learns a joint representation o…”From this paper · §Introduction
- VL-BERT2019 · cited 2דWhile a variety of approaches have been proposed, a large portion of them require object detection for image region feature regression or tagging as part of the pre-training objectives (Tan & Bansal 2019; Su et al. 2020;…”From this paper · §Related Work
- VILLA2020 · cited 2דWhile a variety of approaches have been proposed, a large portion of them require object detection for image region feature regression or tagging as part of the pre-training objectives (Tan & Bansal 2019; Su et al. 2020;…”From this paper · §Related Work
- ViLT2021 · cited 2דSome recent efforts have also explored VLP without object detection module (Xu et al. 2021; Kim et al. 2021; Huang et al. 2021), but they only use clean pretraining data with small scales and thus their zero-shot capabil…”From this paper · §Related Work
- DALL·E2021 · cited 2דMotivated by recent works (Radford et al. 2021; Ramesh et al. 2021; Jia et al. 2021; Tsimpoukelli et al. 2021) that illustrate zero-shot learning in certain image-text tasks, we train our model using large-scale weakly l…”From this paper · §Related Work
Led to
- VLMo2021 · cited 2דThe second category models the interaction of images and text using a deep fusion encoder with cross-modal attention [43, 30, 41, 24, 51, 3, 26, 25, 14, 49, 16, 17, 20, 23, 46].”From VLMo · §Related Work
- METER2021 · cited 1×, 1 in Method“Recently, VL-T5 cho2021unifying and SimVLM wang2021simvlm, on the other hand, advocate the use of a transformer encoder-decoder architecture, where the cross-modal representations are first fed into a decoder and then to…”From METER · §The Meter Framework
- Florence2021 · cited 2×, 1 in Method“Compared with SimVLM (Wang et al. 2021), which uses 1.81.8B image-text pairs, we only use 900900M data to pre-train the image encoder and 2020M for VLP, but achieve better results.”From Florence · §Experiments
- ViTCAP2021 · cited 3דRecent studies have witnessed its great development which are primarily reflected in the aspects of more advanced cross-modal fusion architectures you2016image; vinyals2015show; xu2015show; rennie2017self; yang2019auto;…”From ViTCAP · §Introduction
- BLIP2022 · cited 8×, 1 in Method“Due to the prohibitive expense of acquiring human-annotated texts, most methods (Chen et al. 2020; Li et al. 2020; Li et al. 2021a; Wang et al. 2021; Radford et al. 2021) use image and alt-text pairs crawled from the web…”From BLIP · §Related Work
- Flamingo2022 · cited 2דConcurrent works [124, 17, 119, 154, 58] also propose to formulate numerous vision tasks as text generation problems.”From Flamingo · §Related work
- CoCa2022 · cited 9×, 1 in Method“Unlike other fusion-based foundation methods [16, 35, 17], CoCa is naturally applicable to crossmodal alignment tasks since it generates aligned image and text unimodal embeddings.”From CoCa · §Experiments
- VL-BEiT2022 · cited 3דVision-language pretraining [37, 24, 35, 46, 29, 21, 16, 20, 42, 41, 40, 1, 45] aims to learn multimodal representations from large-scale image-text pairs.”From VL-BEiT · §Related Work
Abstract
With recent progress in joint modeling of visual and textual representations, Vision-Language Pretraining (VLP) has achieved impressive performance on many multimodal downstream tasks. However, the requirement for expensive annotations including clean image captions and regional labels limits the scalability of existing approaches, and complicates the pretraining procedure with the introduction of multiple dataset-specific objectives. In this work, we relax these constraints and present a minimalist pretraining framework, named Simple Visual Language Model (SimVLM). Unlike prior work, SimVLM reduces the training complexity by exploiting large-scale weak supervision, and is trained end-to-end with a single prefix language modeling objective. Without utilizing extra data or task-specific customization, the resulting model significantly outperforms previous pretraining methods and achieves new state-of-the-art results on a wide range of discriminative and generative vision-language benchmarks, including VQA (+3.74% vqa-score), NLVR2 (+1.17% accuracy), SNLI-VE (+1.37% accuracy) and image captioning tasks (+10.1% average CIDEr score). Furthermore, we demonstrate that SimVLM acquires strong generalization and transfer ability, enabling zero-shot behavior including open-ended visual question answering and cross-modality transfer.