VL-BEiT: Generative Vision-Language Pretraining
We introduce a vision-language foundation model called VL-BEiT, which is a bidirectional multimodal Transformer learned by generative pretraining. Our minimalist solution conducts masked prediction on both monomodal and multimodal data with a shared Transformer.
Also cited · not yet reviewed (19)
- VLMo2021 · cited 8×, 1 in Method“Given the image and text representations of monomodal data, and the representations of image-text pairs, we employ a mixture-of-modality-experts (MoME) Transformer [41] to encode different modalities.”From this paper · §Methods
- BEiT2021 · cited 6×, 1 in Method“Following BEiT [2], we apply block-wise masking strategy to mask 40% of image patches.”From this paper · §Methods
- BERT2018 · cited 3×, 1 in Method“Following BERT [7], we randomly mask 15% tokens of monomodal text data.”From this paper · §Methods
- ViT2020 · cited 3×, 1 in Method“Following [9], we split the image 𝒗∈ℝH×W×C{\bm{v}}\in\mathbb{R}^{H\times W\times C} into a sequence of patches, so that the image can be encoded by standard Transformer.”From this paper · §Methods
Show 15 more
- BEiT v22022 · cited 1×, 1 in Method“We use image tokenizer of BEiTv2 [26] to obtain the discrete tokens as the reconstructed targets.”From this paper · §Methods
- ViLT2021 · cited 6דFollowing previous work [16, 41], we use VQA 2.0 dataset [10], and formulate the task as a classification problem to choose the answer from 3,1293,129 most frequent answers.”From this paper · §Experiments
- LXMERT2019 · cited 4דVision-language pretraining [37, 24, 35, 46, 29, 21, 16, 20, 42, 41, 40, 1, 45] aims to learn multimodal representations from large-scale image-text pairs.”From this paper · §Related Work
- CLIP2021 · cited 4דDual-encoder model [29, 14] consists of an image encoder and a text encoder.”From this paper · §Related Work
- MS COCO2014 · cited 3דFor monomodal data, we use ImageNet-22K as the image data, English Wikipedia and BookCorpus [48] as the text data.”From this paper · §Experiments
- VL-BERT2019 · cited 3דPrevious models [24, 35, 21, 46] use an off-the-shelf object detector like Faster R-CNN [30] to obtain image region features.”From this paper · §Related Work
- Oscar2020 · cited 3דPrevious models [24, 35, 21, 46] use an off-the-shelf object detector like Faster R-CNN [30] to obtain image region features.”From this paper · §Related Work
- VinVL2021 · cited 3דPrevious models [24, 35, 21, 46] use an off-the-shelf object detector like Faster R-CNN [30] to obtain image region features.”From this paper · §Related Work
- ALBEF2021 · cited 3דVision-language pretraining [37, 24, 35, 46, 29, 21, 16, 20, 42, 41, 40, 1, 45] aims to learn multimodal representations from large-scale image-text pairs.”From this paper · §Related Work
- SimVLM2021 · cited 3דVision-language pretraining [37, 24, 35, 46, 29, 21, 16, 20, 42, 41, 40, 1, 45] aims to learn multimodal representations from large-scale image-text pairs.”From this paper · §Related Work
- Flickr30k Entities2015 · cited 2דWe conduct vision-language finetuning experiments on the widely used visual question answering [10], natural language for visual reasoning [36] and image-text retrieval [27, 22] tasks.”From this paper · §Experiments
- VQA v22016 · cited 2דFollowing previous work [16, 41], we use VQA 2.0 dataset [10], and formulate the task as a classification problem to choose the answer from 3,1293,129 most frequent answers.”From this paper · §Experiments
- NLVR22018 · cited 2דWe use NLVR2 [36] dataset to evaluate the model.”From this paper · §Experiments
- ViLBERT2019 · cited 2דPrevious models [24, 35, 21, 46] use an off-the-shelf object detector like Faster R-CNN [30] to obtain image region features.”From this paper · §Related Work
- ALIGN2021 · cited 2דDual-encoder model [29, 14] consists of an image encoder and a text encoder.”From this paper · §Related Work
Led to
- BEiT-32022 · cited 2×, 1 in Method“Second, the pretraining task based on masked data modeling has been successfully applied to various modalities, such as texts [14], images [3, 40], and image-text pairs [7].”From BEiT-3 · §Introduction: The Big Convergence
Abstract
We introduce a vision-language foundation model called VL-BEiT, which is a bidirectional multimodal Transformer learned by generative pretraining. Our minimalist solution conducts masked prediction on both monomodal and multimodal data with a shared Transformer. Specifically, we perform masked vision-language modeling on image-text pairs, masked language modeling on texts, and masked image modeling on images. VL-BEiT is learned from scratch with one unified pretraining task, one shared backbone, and one-stage training. Our method is conceptually simple and empirically effective. Experimental results show that VL-BEiT obtains strong results on various vision-language benchmarks, such as visual question answering, visual reasoning, and image-text retrieval. Moreover, our method learns transferable visual features, achieving competitive performance on image classification, and semantic segmentation.