VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts
We present a unified Vision-Language pretrained Model (VLMo) that jointly learns a dual encoder and a fusion encoder with a modular Transformer network. Specifically, we introduce Mixture-of-Modality-Experts (MoME) Transformer, where each block contains a pool of modality-specific experts and a shared self-attention layer.
Also cited · not yet reviewed (24)
- BERT2018 · cited 6×, 3 in Method“Following BERT [10], we tokenize the text to subword units by WordPiece [47].”From this paper · §Methods
- ALBEF2021 · cited 9×, 2 in Method“Pre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”From this paper · §Related Work
- BEiT2021 · cited 6×, 2 in Method“Following vision Transformers [12, 44, 2], the 2D image 𝒗∈ℝH×W×C{\bm{v}}\in\mathbb{R}^{H\times W\times C} is split and reshaped into N=HW/P2N={HW}/{P^{2}} patches 𝒗p∈ℝN×(P2C){\bm{v}}^{p}\in\mathbb{R}^{N\times(P^{2}C)…”From this paper · §Methods
- ViT2020 · cited 5×, 1 in Method“Following vision Transformers [12, 44, 2], the 2D image 𝒗∈ℝH×W×C{\bm{v}}\in\mathbb{R}^{H\times W\times C} is split and reshaped into N=HW/P2N={HW}/{P^{2}} patches 𝒗p∈ℝN×(P2C){\bm{v}}^{p}\in\mathbb{R}^{N\times(P^{2}C)…”From this paper · §Methods
Show 20 more
- DeiT2020 · cited 3×, 1 in Method“Following vision Transformers [12, 44, 2], the 2D image 𝒗∈ℝH×W×C{\bm{v}}\in\mathbb{R}^{H\times W\times C} is split and reshaped into N=HW/P2N={HW}/{P^{2}} patches 𝒗p∈ℝN×(P2C){\bm{v}}^{p}\in\mathbb{R}^{N\times(P^{2}C)…”From this paper · §Methods
- GNMT2016 · cited 1×, 1 in Method“Following BERT [10], we tokenize the text to subword units by WordPiece [47].”From this paper · §Methods
- Switch Transformer2021 · cited 1×, 1 in Method“Inspired by mixture-of-experts networks [40, 13], we propose a general-purpose multimodal Transformer for vision-language tasks, namely MoME Transformer, to encode different modalities.”From this paper · §Methods
- ViLT2021 · cited 10דPre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”From this paper · §Related Work
- UNITER2019 · cited 6דPre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”From this paper · §Related Work
- VL-BERT2019 · cited 5דPre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”From this paper · §Related Work
- ViLBERT2019 · cited 4דThis leads to a significant speedup compared with previous models using image region features, which are extracted by an off-the-shelf object detector [30, 41, 3, 14, 25, 49].”From this paper · §Experiments
- VinVL2021 · cited 4דPre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”From this paper · §Related Work
- CLIP2021 · cited 4דPre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”From this paper · §Related Work
- LXMERT2019 · cited 3דPre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”From this paper · §Related Work
- Oscar2020 · cited 3דFollowing OSCAR [26] and VinVL [49], we convert the triplet input to two image-text pairs, each containing the text description and one image.”From this paper · §Experiments
- VILLA2020 · cited 3דThis leads to a significant speedup compared with previous models using image region features, which are extracted by an off-the-shelf object detector [30, 41, 3, 14, 25, 49].”From this paper · §Experiments
- ALIGN2021 · cited 3דPre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”From this paper · §Related Work
- MS COCO2014 · cited 2דFollowing previous work [3, 20], our pre-training data consists of four image captioning datasets: Conceptual Captions (CC) [39], SBU Captions [32], COCO [27] and Visual Genome (VG) [21] datasets.”From this paper · §Experiments
- VQA v22016 · cited 2דWe train and evaluate the model on VQA 2.0 dataset [15].”From this paper · §Experiments
- Transformer2017 · cited 2דPre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”From this paper · §Related Work
- NLVR22018 · cited 2דWe first conduct fine-tuning experiments on two widely used classification datasets: visual question answering [15] and natural language for visual reasoning [42].”From this paper · §Experiments
- UniLM2019 · cited 2דPre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”From this paper · §Related Work
- SimVLM2021 · cited 2דThe second category models the interaction of images and text using a deep fusion encoder with cross-modal attention [43, 30, 41, 24, 51, 3, 26, 25, 14, 49, 16, 17, 20, 23, 46].”From this paper · §Related Work
- Florence2021 · cited 2דOur large-size model even outperforms SimVLM-Huge [46] and Florence-Huge [48] by a large margin, which consists of more parameters and are also trained on larger-scale image-text pairs.”From this paper · §Experiments
Led to
- CoCa2022 · cited 2×, 1 in Method“On the other hand, while many existing methods [32, 33, 35, 30, 36, 37] train model components with multiple stages on various data sources and/or modalities, CoCa is pretrained end-to-end from scratch directly with vari…”From CoCa · §Approach
- VL-BEiT2022 · cited 8×, 1 in Method“Given the image and text representations of monomodal data, and the representations of image-text pairs, we employ a mixture-of-modality-experts (MoME) Transformer [41] to encode different modalities.”From VL-BEiT · §Methods
- BEiT-32022 · cited 5×, 2 in Method“We use Multiway Transformers [53] as the backbone model to encode different modalities.”From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
Abstract
We present a unified Vision-Language pretrained Model (VLMo) that jointly learns a dual encoder and a fusion encoder with a modular Transformer network. Specifically, we introduce Mixture-of-Modality-Experts (MoME) Transformer, where each block contains a pool of modality-specific experts and a shared self-attention layer. Because of the modeling flexibility of MoME, pretrained VLMo can be fine-tuned as a fusion encoder for vision-language classification tasks, or used as a dual encoder for efficient image-text retrieval. Moreover, we propose a stagewise pre-training strategy, which effectively leverages large-scale image-only and text-only data besides image-text pairs. Experimental results show that VLMo achieves state-of-the-art results on various vision-language tasks, including VQA, NLVR2 and image-text retrieval. The code and pretrained models are available at https://aka.ms/vlmo.