Paper Lineage
Esc
MethodJun 2022arXiv 2206.01127cs.CV

VL-BEiT: Generative Vision-Language Pretraining

Hangbo Bao, Wenhui Wang, Li Dong, Furu Wei

We introduce a vision-language foundation model called VL-BEiT, which is a bidirectional multimodal Transformer learned by generative pretraining. Our minimalist solution conducts masked prediction on both monomodal and multimodal data with a shared Transformer.

From the abstract

Built on

19 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (19)

  • VLMo2021 · cited 8×, 1 in Method
    “Given the image and text representations of monomodal data, and the representations of image-text pairs, we employ a mixture-of-modality-experts (MoME) Transformer [41] to encode different modalities.”
    From this paper · §Methods
  • BEiT2021 · cited 6×, 1 in Method
    “Following BEiT [2], we apply block-wise masking strategy to mask 40% of image patches.”
    From this paper · §Methods
  • BERT2018 · cited 3×, 1 in Method
    “Following BERT [7], we randomly mask 15% tokens of monomodal text data.”
    From this paper · §Methods
  • ViT2020 · cited 3×, 1 in Method
    “Following [9], we split the image 𝒗∈ℝH×W×C{\bm{v}}\in\mathbb{R}^{H\times W\times C} into a sequence of patches, so that the image can be encoded by standard Transformer.”
    From this paper · §Methods
Show 15 more
  • BEiT v22022 · cited 1×, 1 in Method
    “We use image tokenizer of BEiTv2 [26] to obtain the discrete tokens as the reconstructed targets.”
    From this paper · §Methods
  • ViLT2021 · cited 6×
    “Following previous work [16, 41], we use VQA 2.0 dataset [10], and formulate the task as a classification problem to choose the answer from 3,1293,129 most frequent answers.”
    From this paper · §Experiments
  • LXMERT2019 · cited 4×
    “Vision-language pretraining [37, 24, 35, 46, 29, 21, 16, 20, 42, 41, 40, 1, 45] aims to learn multimodal representations from large-scale image-text pairs.”
    From this paper · §Related Work
  • CLIP2021 · cited 4×
    “Dual-encoder model [29, 14] consists of an image encoder and a text encoder.”
    From this paper · §Related Work
  • MS COCO2014 · cited 3×
    “For monomodal data, we use ImageNet-22K as the image data, English Wikipedia and BookCorpus [48] as the text data.”
    From this paper · §Experiments
  • VL-BERT2019 · cited 3×
    “Previous models [24, 35, 21, 46] use an off-the-shelf object detector like Faster R-CNN [30] to obtain image region features.”
    From this paper · §Related Work
  • Oscar2020 · cited 3×
    “Previous models [24, 35, 21, 46] use an off-the-shelf object detector like Faster R-CNN [30] to obtain image region features.”
    From this paper · §Related Work
  • VinVL2021 · cited 3×
    “Previous models [24, 35, 21, 46] use an off-the-shelf object detector like Faster R-CNN [30] to obtain image region features.”
    From this paper · §Related Work
  • ALBEF2021 · cited 3×
    “Vision-language pretraining [37, 24, 35, 46, 29, 21, 16, 20, 42, 41, 40, 1, 45] aims to learn multimodal representations from large-scale image-text pairs.”
    From this paper · §Related Work
  • SimVLM2021 · cited 3×
    “Vision-language pretraining [37, 24, 35, 46, 29, 21, 16, 20, 42, 41, 40, 1, 45] aims to learn multimodal representations from large-scale image-text pairs.”
    From this paper · §Related Work
  • Flickr30k Entities2015 · cited 2×
    “We conduct vision-language finetuning experiments on the widely used visual question answering [10], natural language for visual reasoning [36] and image-text retrieval [27, 22] tasks.”
    From this paper · §Experiments
  • VQA v22016 · cited 2×
    “Following previous work [16, 41], we use VQA 2.0 dataset [10], and formulate the task as a classification problem to choose the answer from 3,1293,129 most frequent answers.”
    From this paper · §Experiments
  • NLVR22018 · cited 2×
    “We use NLVR2 [36] dataset to evaluate the model.”
    From this paper · §Experiments
  • ViLBERT2019 · cited 2×
    “Previous models [24, 35, 21, 46] use an off-the-shelf object detector like Faster R-CNN [30] to obtain image region features.”
    From this paper · §Related Work
  • ALIGN2021 · cited 2×
    “Dual-encoder model [29, 14] consists of an image encoder and a text encoder.”
    From this paper · §Related Work

Led to

  • BEiT-32022 · cited 2×, 1 in Method
    “Second, the pretraining task based on masked data modeling has been successfully applied to various modalities, such as texts [14], images [3, 40], and image-text pairs [7].”
    From BEiT-3 · §Introduction: The Big Convergence
Abstract

We introduce a vision-language foundation model called VL-BEiT, which is a bidirectional multimodal Transformer learned by generative pretraining. Our minimalist solution conducts masked prediction on both monomodal and multimodal data with a shared Transformer. Specifically, we perform masked vision-language modeling on image-text pairs, masked language modeling on texts, and masked image modeling on images. VL-BEiT is learned from scratch with one unified pretraining task, one shared backbone, and one-stage training. Our method is conceptually simple and empirically effective. Experimental results show that VL-BEiT obtains strong results on various vision-language benchmarks, such as visual question answering, visual reasoning, and image-text retrieval. Moreover, our method learns transferable visual features, achieving competitive performance on image classification, and semantic segmentation.