Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks
A big convergence of language, vision, and multimodal pretraining is emerging. In this work, we introduce a general-purpose multimodal foundation model BEiT-3, which achieves state-of-the-art transfer performance on both vision and vision-language tasks.
Also cited · not yet reviewed (29)
- CoCa2022 · cited 6×, 4 in Method“In contrast, previous vision-language models [36, 65, 26, 34, 53, 30, 62] usually employ multiple pretraining tasks, such as image-text contrast, image-text matching, and word-patch/region alignment.”From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
- BEiT2021 · cited 7×, 3 in Method“Second, the pretraining task based on masked data modeling has been successfully applied to various modalities, such as texts [14], images [3, 40], and image-text pairs [7].”From this paper · §Introduction: The Big Convergence
- CLIP2021 · cited 5×, 3 in Method“In comparison, contrastive-based models [43, 23, 60, 62] usually need a very large batch size11 1 For example, CoCa [62] uses 6565k batch size, CLIP [43] uses 3232k batch size, and Florence [60] uses 2424k batch size.”From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
- VLMo2021 · cited 5×, 2 in Method“We use Multiway Transformers [53] as the backbone model to encode different modalities.”From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
Show 25 more
- BEiT v22022 · cited 4×, 2 in Method“Second, the pretraining task based on masked data modeling has been successfully applied to various modalities, such as texts [14], images [3, 40], and image-text pairs [7].”From this paper · §Introduction: The Big Convergence
- Florence2021 · cited 3×, 2 in Method“In comparison, contrastive-based models [43, 23, 60, 62] usually need a very large batch size11 1 For example, CoCa [62] uses 6565k batch size, CLIP [43] uses 3232k batch size, and Florence [60] uses 2424k batch size.”From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
- SentencePiece2018 · cited 2×, 2 in Method“Specifically, text data is tokenized by a SentencePiece tokenizer [25].”From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
- ALIGN2021 · cited 2×, 2 in Method“In comparison, contrastive-based models [43, 23, 60, 62] usually need a very large batch size11 1 For example, CoCa [62] uses 6565k batch size, CLIP [43] uses 3232k batch size, and Florence [60] uses 2424k batch size.”From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
- MS COCO2014 · cited 6×, 1 in Method“For multimodal data, there are about 1515M images and 2121M image-text pairs collected from five public datasets: Conceptual 12M (CC12M) [11], Conceptual Captions (CC3M) [46], SBU Captions (SBU) [39], COCO [31] and Visua…”From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
- ViLT2021 · cited 6×, 1 in Method“In contrast, previous vision-language models [36, 65, 26, 34, 53, 30, 62] usually employ multiple pretraining tasks, such as image-text contrast, image-text matching, and word-patch/region alignment.”From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
- VinVL2021 · cited 4×, 1 in Method“In contrast, previous vision-language models [36, 65, 26, 34, 53, 30, 62] usually employ multiple pretraining tasks, such as image-text contrast, image-text matching, and word-patch/region alignment.”From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
- Scaling ViTs (ViT-G)2021 · cited 2×, 1 in Method“BEiT-3 is a giant-size foundation model following the setup of ViT-giant [63].”From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
- VL-BEiT2022 · cited 2×, 1 in Method“Second, the pretraining task based on masked data modeling has been successfully applied to various modalities, such as texts [14], images [3, 40], and image-text pairs [7].”From this paper · §Introduction: The Big Convergence
- BookCorpus (books & movies)2015 · cited 1×, 1 in Method“For monomodal data, we use 1414M images from ImageNet-21K and 160160GB text corpora [4] from English Wikipedia, BookCorpus [64], OpenWebText22 2 http://skylion007.github.io/OpenWebTextCorpus, CC-News [33], and Stories [5…”From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
- Visual Genome2016 · cited 1×, 1 in Method“For multimodal data, there are about 1515M images and 2121M image-text pairs collected from five public datasets: Conceptual 12M (CC12M) [11], Conceptual Captions (CC3M) [46], SBU Captions (SBU) [39], COCO [31] and Visua…”From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
- AdamW2017 · cited 1×, 1 in Method“We use the AdamW [28] optimizer with β1=0.9\beta_{1}=0.9, β2=0.98\beta_{2}=0.98 and ϵ=\epsilon=1e-6 for optimization.”From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
- Instance discrimination2018 · cited 1×, 1 in Method“We use the same image augmentation as in BEiT [3], including random resized cropping, horizontal flipping, and color jittering [56].”From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
- RoBERTa2019 · cited 1×, 1 in Method“For monomodal data, we use 1414M images from ImageNet-21K and 160160GB text corpora [4] from English Wikipedia, BookCorpus [64], OpenWebText22 2 http://skylion007.github.io/OpenWebTextCorpus, CC-News [33], and Stories [5…”From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
- Oscar2020 · cited 1×, 1 in Method“In contrast, previous vision-language models [36, 65, 26, 34, 53, 30, 62] usually employ multiple pretraining tasks, such as image-text contrast, image-text matching, and word-patch/region alignment.”From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
- Conceptual 12M2021 · cited 1×, 1 in Method“For multimodal data, there are about 1515M images and 2121M image-text pairs collected from five public datasets: Conceptual 12M (CC12M) [11], Conceptual Captions (CC3M) [46], SBU Captions (SBU) [39], COCO [31] and Visua…”From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
- ALBEF2021 · cited 1×, 1 in Method“In contrast, previous vision-language models [36, 65, 26, 34, 53, 30, 62] usually employ multiple pretraining tasks, such as image-text contrast, image-text matching, and word-patch/region alignment.”From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
- BLIP2022 · cited 1×, 1 in Method“In contrast, previous vision-language models [36, 65, 26, 34, 53, 30, 62] usually employ multiple pretraining tasks, such as image-text contrast, image-text matching, and word-patch/region alignment.”From this paper · §BEiT-3: A General-Purpose Multimodal Foundation Model
- BERT2018 · cited 3דSecond, the pretraining task based on masked data modeling has been successfully applied to various modalities, such as texts [14], images [3, 40], and image-text pairs [7].”From this paper · §Introduction: The Big Convergence
- Karpathy visual-semantic alignment2014 · cited 2דWe use the COCO [31] benchmark, finetune and evaluate the model on Karpathy split [24].”From this paper · §Experiments on Vision and Vision-Language Tasks
- Flickr30k Entities2015 · cited 2דWe evaluate the capabilities of BEiT-3 on the widely used vision-language understanding and generation benchmarks, including visual question answering [19], visual reasoning [49], image-text retrieval [41, 31], and image…”From this paper · §Experiments on Vision and Vision-Language Tasks
- VQA v22016 · cited 2דFollowing previous work [2, 65, 26], we conduct finetuning experiments on the VQA v2.0 dataset [19] and formulate the task as a classification problem.”From this paper · §Experiments on Vision and Vision-Language Tasks
- NLVR22018 · cited 2דWe evaluate the capabilities of BEiT-3 on the widely used vision-language understanding and generation benchmarks, including visual question answering [19], visual reasoning [49], image-text retrieval [41, 31], and image…”From this paper · §Experiments on Vision and Vision-Language Tasks
- UniLM2019 · cited 2דFollowing UniLM [17] and s2s-ft [5], BEiT-3 is used as a conditional generation model via masked finetuning.”From this paper · §Experiments on Vision and Vision-Language Tasks
- ViT2020 · cited 2דRather than appending a task layer to the vision encoder [13, 3], we formulate the task as an image-to-text retrieval task.”From this paper · §Experiments on Vision and Vision-Language Tasks
Led to
- EVA2022 · cited 11דHowever, there remains a debate that (i) tokenized semantic features could provide better supervision signal for masked modeling in vision bao2021beit; beitv2; beit3, and (ii) good performances could be also achieved via…”From EVA · §Introduction
- BLIP-22023 · cited 3דDepending on the downstream task, different model architectures have been proposed, including the dual-encoder architecture (Radford et al. 2021; Jia et al. 2021), the fusion-encoder architecture (Tan & Bansal 2019; Li e…”From BLIP-2 · §Related Work
Abstract
A big convergence of language, vision, and multimodal pretraining is emerging. In this work, we introduce a general-purpose multimodal foundation model BEiT-3, which achieves state-of-the-art transfer performance on both vision and vision-language tasks. Specifically, we advance the big convergence from three aspects: backbone architecture, pretraining task, and model scaling up. We introduce Multiway Transformers for general-purpose modeling, where the modular architecture enables both deep fusion and modality-specific encoding. Based on the shared backbone, we perform masked "language" modeling on images (Imglish), texts (English), and image-text pairs ("parallel sentences") in a unified manner. Experimental results show that BEiT-3 obtains state-of-the-art performance on object detection (COCO), semantic segmentation (ADE20K), image classification (ImageNet), visual reasoning (NLVR2), visual question answering (VQAv2), image captioning (COCO), and cross-modal retrieval (Flickr30K, COCO).