Paper Lineage
Esc
MethodNov 2021arXiv 2111.02358cs.CV

VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts

Hangbo Bao, Wenhui Wang, Li Dong and 5 others

We present a unified Vision-Language pretrained Model (VLMo) that jointly learns a dual encoder and a fusion encoder with a modular Transformer network. Specifically, we introduce Mixture-of-Modality-Experts (MoME) Transformer, where each block contains a pool of modality-specific experts and a shared self-attention layer.

From the abstract

Built on

24 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (24)

  • BERT2018 · cited 6×, 3 in Method
    “Following BERT [10], we tokenize the text to subword units by WordPiece [47].”
    From this paper · §Methods
  • ALBEF2021 · cited 9×, 2 in Method
    “Pre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”
    From this paper · §Related Work
  • BEiT2021 · cited 6×, 2 in Method
    “Following vision Transformers [12, 44, 2], the 2D image 𝒗∈ℝH×W×C{\bm{v}}\in\mathbb{R}^{H\times W\times C} is split and reshaped into N=H​W/P2N={HW}/{P^{2}} patches 𝒗p∈ℝN×(P2​C){\bm{v}}^{p}\in\mathbb{R}^{N\times(P^{2}C)…”
    From this paper · §Methods
  • ViT2020 · cited 5×, 1 in Method
    “Following vision Transformers [12, 44, 2], the 2D image 𝒗∈ℝH×W×C{\bm{v}}\in\mathbb{R}^{H\times W\times C} is split and reshaped into N=H​W/P2N={HW}/{P^{2}} patches 𝒗p∈ℝN×(P2​C){\bm{v}}^{p}\in\mathbb{R}^{N\times(P^{2}C)…”
    From this paper · §Methods
Show 20 more
  • DeiT2020 · cited 3×, 1 in Method
    “Following vision Transformers [12, 44, 2], the 2D image 𝒗∈ℝH×W×C{\bm{v}}\in\mathbb{R}^{H\times W\times C} is split and reshaped into N=H​W/P2N={HW}/{P^{2}} patches 𝒗p∈ℝN×(P2​C){\bm{v}}^{p}\in\mathbb{R}^{N\times(P^{2}C)…”
    From this paper · §Methods
  • GNMT2016 · cited 1×, 1 in Method
    “Following BERT [10], we tokenize the text to subword units by WordPiece [47].”
    From this paper · §Methods
  • Switch Transformer2021 · cited 1×, 1 in Method
    “Inspired by mixture-of-experts networks [40, 13], we propose a general-purpose multimodal Transformer for vision-language tasks, namely MoME Transformer, to encode different modalities.”
    From this paper · §Methods
  • ViLT2021 · cited 10×
    “Pre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”
    From this paper · §Related Work
  • UNITER2019 · cited 6×
    “Pre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”
    From this paper · §Related Work
  • VL-BERT2019 · cited 5×
    “Pre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”
    From this paper · §Related Work
  • ViLBERT2019 · cited 4×
    “This leads to a significant speedup compared with previous models using image region features, which are extracted by an off-the-shelf object detector [30, 41, 3, 14, 25, 49].”
    From this paper · §Experiments
  • VinVL2021 · cited 4×
    “Pre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”
    From this paper · §Related Work
  • CLIP2021 · cited 4×
    “Pre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”
    From this paper · §Related Work
  • LXMERT2019 · cited 3×
    “Pre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”
    From this paper · §Related Work
  • Oscar2020 · cited 3×
    “Following OSCAR [26] and VinVL [49], we convert the triplet input to two image-text pairs, each containing the text description and one image.”
    From this paper · §Experiments
  • VILLA2020 · cited 3×
    “This leads to a significant speedup compared with previous models using image region features, which are extracted by an off-the-shelf object detector [30, 41, 3, 14, 25, 49].”
    From this paper · §Experiments
  • ALIGN2021 · cited 3×
    “Pre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”
    From this paper · §Related Work
  • MS COCO2014 · cited 2×
    “Following previous work [3, 20], our pre-training data consists of four image captioning datasets: Conceptual Captions (CC) [39], SBU Captions [32], COCO [27] and Visual Genome (VG) [21] datasets.”
    From this paper · §Experiments
  • VQA v22016 · cited 2×
    “We train and evaluate the model on VQA 2.0 dataset [15].”
    From this paper · §Experiments
  • Transformer2017 · cited 2×
    “Pre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”
    From this paper · §Related Work
  • NLVR22018 · cited 2×
    “We first conduct fine-tuning experiments on two widely used classification datasets: visual question answering [15] and natural language for visual reasoning [42].”
    From this paper · §Experiments
  • UniLM2019 · cited 2×
    “Pre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”
    From this paper · §Related Work
  • SimVLM2021 · cited 2×
    “The second category models the interaction of images and text using a deep fusion encoder with cross-modal attention [43, 30, 41, 24, 51, 3, 26, 25, 14, 49, 16, 17, 20, 23, 46].”
    From this paper · §Related Work
  • Florence2021 · cited 2×
    “Our large-size model even outperforms SimVLM-Huge [46] and Florence-Huge [48] by a large margin, which consists of more parameters and are also trained on larger-scale image-text pairs.”
    From this paper · §Experiments

Led to

  • CoCa2022 · cited 2×, 1 in Method
    “On the other hand, while many existing methods [32, 33, 35, 30, 36, 37] train model components with multiple stages on various data sources and/or modalities, CoCa is pretrained end-to-end from scratch directly with vari…”
    From CoCa · §Approach
  • VL-BEiT2022 · cited 8×, 1 in Method
    “Given the image and text representations of monomodal data, and the representations of image-text pairs, we employ a mixture-of-modality-experts (MoME) Transformer [41] to encode different modalities.”
    From VL-BEiT · §Methods
  • BEiT-32022 · cited 5×, 2 in Method
    “We use Multiway Transformers [53] as the backbone model to encode different modalities.”
    From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
Abstract

We present a unified Vision-Language pretrained Model (VLMo) that jointly learns a dual encoder and a fusion encoder with a modular Transformer network. Specifically, we introduce Mixture-of-Modality-Experts (MoME) Transformer, where each block contains a pool of modality-specific experts and a shared self-attention layer. Because of the modeling flexibility of MoME, pretrained VLMo can be fine-tuned as a fusion encoder for vision-language classification tasks, or used as a dual encoder for efficient image-text retrieval. Moreover, we propose a stagewise pre-training strategy, which effectively leverages large-scale image-only and text-only data besides image-text pairs. Experimental results show that VLMo achieves state-of-the-art results on various vision-language tasks, including VQA, NLVR2 and image-text retrieval. The code and pretrained models are available at https://aka.ms/vlmo.