Paper Lineage
Esc
MethodAug 2022arXiv 2208.10442cs.CV

Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Wenhui Wang, Hangbo Bao, Li Dong and 8 others

A big convergence of language, vision, and multimodal pretraining is emerging. In this work, we introduce a general-purpose multimodal foundation model BEiT-3, which achieves state-of-the-art transfer performance on both vision and vision-language tasks.

From the abstract

Built on

29 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (29)

  • CoCa2022 · cited 6×, 4 in Method
    “In contrast, previous vision-language models [36, 65, 26, 34, 53, 30, 62] usually employ multiple pretraining tasks, such as image-text contrast, image-text matching, and word-patch/region alignment.”
    From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • BEiT2021 · cited 7×, 3 in Method
    “Second, the pretraining task based on masked data modeling has been successfully applied to various modalities, such as texts [14], images [3, 40], and image-text pairs [7].”
    From this paper · §Introduction: The Big Convergence
  • CLIP2021 · cited 5×, 3 in Method
    “In comparison, contrastive-based models [43, 23, 60, 62] usually need a very large batch size11 1 For example, CoCa [62] uses 6565k batch size, CLIP [43] uses 3232k batch size, and Florence [60] uses 2424k batch size.”
    From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • VLMo2021 · cited 5×, 2 in Method
    “We use Multiway Transformers [53] as the backbone model to encode different modalities.”
    From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
Show 25 more
  • BEiT v22022 · cited 4×, 2 in Method
    “Second, the pretraining task based on masked data modeling has been successfully applied to various modalities, such as texts [14], images [3, 40], and image-text pairs [7].”
    From this paper · §Introduction: The Big Convergence
  • Florence2021 · cited 3×, 2 in Method
    “In comparison, contrastive-based models [43, 23, 60, 62] usually need a very large batch size11 1 For example, CoCa [62] uses 6565k batch size, CLIP [43] uses 3232k batch size, and Florence [60] uses 2424k batch size.”
    From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • SentencePiece2018 · cited 2×, 2 in Method
    “Specifically, text data is tokenized by a SentencePiece tokenizer [25].”
    From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • ALIGN2021 · cited 2×, 2 in Method
    “In comparison, contrastive-based models [43, 23, 60, 62] usually need a very large batch size11 1 For example, CoCa [62] uses 6565k batch size, CLIP [43] uses 3232k batch size, and Florence [60] uses 2424k batch size.”
    From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • MS COCO2014 · cited 6×, 1 in Method
    “For multimodal data, there are about 1515M images and 2121M image-text pairs collected from five public datasets: Conceptual 12M (CC12M) [11], Conceptual Captions (CC3M) [46], SBU Captions (SBU) [39], COCO [31] and Visua…”
    From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • ViLT2021 · cited 6×, 1 in Method
    “In contrast, previous vision-language models [36, 65, 26, 34, 53, 30, 62] usually employ multiple pretraining tasks, such as image-text contrast, image-text matching, and word-patch/region alignment.”
    From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • VinVL2021 · cited 4×, 1 in Method
    “In contrast, previous vision-language models [36, 65, 26, 34, 53, 30, 62] usually employ multiple pretraining tasks, such as image-text contrast, image-text matching, and word-patch/region alignment.”
    From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • Scaling ViTs (ViT-G)2021 · cited 2×, 1 in Method
    “BEiT-3 is a giant-size foundation model following the setup of ViT-giant [63].”
    From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • VL-BEiT2022 · cited 2×, 1 in Method
    “Second, the pretraining task based on masked data modeling has been successfully applied to various modalities, such as texts [14], images [3, 40], and image-text pairs [7].”
    From this paper · §Introduction: The Big Convergence
  • BookCorpus (books & movies)2015 · cited 1×, 1 in Method
    “For monomodal data, we use 1414M images from ImageNet-21K and 160160GB text corpora [4] from English Wikipedia, BookCorpus [64], OpenWebText22 2 http://skylion007.github.io/OpenWebTextCorpus, CC-News [33], and Stories [5…”
    From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • Visual Genome2016 · cited 1×, 1 in Method
    “For multimodal data, there are about 1515M images and 2121M image-text pairs collected from five public datasets: Conceptual 12M (CC12M) [11], Conceptual Captions (CC3M) [46], SBU Captions (SBU) [39], COCO [31] and Visua…”
    From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • AdamW2017 · cited 1×, 1 in Method
    “We use the AdamW [28] optimizer with β1=0.9\beta_{1}=0.9, β2=0.98\beta_{2}=0.98 and ϵ=\epsilon=1e-6 for optimization.”
    From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • Instance discrimination2018 · cited 1×, 1 in Method
    “We use the same image augmentation as in BEiT [3], including random resized cropping, horizontal flipping, and color jittering [56].”
    From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • RoBERTa2019 · cited 1×, 1 in Method
    “For monomodal data, we use 1414M images from ImageNet-21K and 160160GB text corpora [4] from English Wikipedia, BookCorpus [64], OpenWebText22 2 http://skylion007.github.io/OpenWebTextCorpus, CC-News [33], and Stories [5…”
    From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • Oscar2020 · cited 1×, 1 in Method
    “In contrast, previous vision-language models [36, 65, 26, 34, 53, 30, 62] usually employ multiple pretraining tasks, such as image-text contrast, image-text matching, and word-patch/region alignment.”
    From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • Conceptual 12M2021 · cited 1×, 1 in Method
    “For multimodal data, there are about 1515M images and 2121M image-text pairs collected from five public datasets: Conceptual 12M (CC12M) [11], Conceptual Captions (CC3M) [46], SBU Captions (SBU) [39], COCO [31] and Visua…”
    From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • ALBEF2021 · cited 1×, 1 in Method
    “In contrast, previous vision-language models [36, 65, 26, 34, 53, 30, 62] usually employ multiple pretraining tasks, such as image-text contrast, image-text matching, and word-patch/region alignment.”
    From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • BLIP2022 · cited 1×, 1 in Method
    “In contrast, previous vision-language models [36, 65, 26, 34, 53, 30, 62] usually employ multiple pretraining tasks, such as image-text contrast, image-text matching, and word-patch/region alignment.”
    From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • BERT2018 · cited 3×
    “Second, the pretraining task based on masked data modeling has been successfully applied to various modalities, such as texts [14], images [3, 40], and image-text pairs [7].”
    From this paper · §Introduction: The Big Convergence
  • “We use the COCO [31] benchmark, finetune and evaluate the model on Karpathy split [24].”
    From this paper · §Experiments on Vision and Vision-Language Tasks
  • Flickr30k Entities2015 · cited 2×
    “We evaluate the capabilities of BEiT-3 on the widely used vision-language understanding and generation benchmarks, including visual question answering [19], visual reasoning [49], image-text retrieval [41, 31], and image…”
    From this paper · §Experiments on Vision and Vision-Language Tasks
  • VQA v22016 · cited 2×
    “Following previous work [2, 65, 26], we conduct finetuning experiments on the VQA v2.0 dataset [19] and formulate the task as a classification problem.”
    From this paper · §Experiments on Vision and Vision-Language Tasks
  • NLVR22018 · cited 2×
    “We evaluate the capabilities of BEiT-3 on the widely used vision-language understanding and generation benchmarks, including visual question answering [19], visual reasoning [49], image-text retrieval [41, 31], and image…”
    From this paper · §Experiments on Vision and Vision-Language Tasks
  • UniLM2019 · cited 2×
    “Following UniLM [17] and s2s-ft [5], BEiT-3 is used as a conditional generation model via masked finetuning.”
    From this paper · §Experiments on Vision and Vision-Language Tasks
  • ViT2020 · cited 2×
    “Rather than appending a task layer to the vision encoder [13, 3], we formulate the task as an image-to-text retrieval task.”
    From this paper · §Experiments on Vision and Vision-Language Tasks

Led to

  • EVA2022 · cited 11×
    “However, there remains a debate that (i) tokenized semantic features could provide better supervision signal for masked modeling in vision bao2021beit; beitv2; beit3, and (ii) good performances could be also achieved via…”
    From EVA · §Introduction
  • BLIP-22023 · cited 3×
    “Depending on the downstream task, different model architectures have been proposed, including the dual-encoder architecture (Radford et al. 2021; Jia et al. 2021), the fusion-encoder architecture (Tan & Bansal 2019; Li e…”
    From BLIP-2 · §Related Work
Abstract

A big convergence of language, vision, and multimodal pretraining is emerging. In this work, we introduce a general-purpose multimodal foundation model BEiT-3, which achieves state-of-the-art transfer performance on both vision and vision-language tasks. Specifically, we advance the big convergence from three aspects: backbone architecture, pretraining task, and model scaling up. We introduce Multiway Transformers for general-purpose modeling, where the modular architecture enables both deep fusion and modality-specific encoding. Based on the shared backbone, we perform masked "language" modeling on images (Imglish), texts (English), and image-text pairs ("parallel sentences") in a unified manner. Experimental results show that BEiT-3 obtains state-of-the-art performance on object detection (COCO), semantic segmentation (ADE20K), image classification (ImageNet), visual reasoning (NLVR2), visual question answering (VQAv2), image captioning (COCO), and cross-modal retrieval (Flickr30K, COCO).