Paper Lineage
Esc
MethodJul 2021arXiv 2107.07651cs.CV

Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Deepak Gotmare and 3 others

Large-scale vision and language representation learning has shown promising improvements on various vision-language tasks. Most existing methods employ a transformer-based multimodal encoder to jointly model visual tokens (region-based image features) and word tokens.

From the abstract

Built on

20 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (20)

  • ViT2020 · cited 2×, 2 in Method
    “We use a 12-layer visual transformer ViT-B/16 [38] as the image encoder, and initialize it with weights pre-trained on ImageNet-1k from [31].”
    From this paper · §ALBEF Pre-training
  • UNITER2019 · cited 7×, 1 in Method
    “Following UNITER [2], we construct our pre-training data using two web datasets (Conceptual Captions [4], SBU Captions [5]) and two in-domain datasets (COCO [41] and Visual Genome [42]).”
    From this paper · §ALBEF Pre-training
  • MoCo2019 · cited 3×, 1 in Method
    “Inspired by MoCo [24], we maintain two queues to store the most recent MM image-text representations from the momentum unimodal encoders.”
    From this paper · §ALBEF Pre-training
  • DeiT2020 · cited 2×, 1 in Method
    “We use a 12-layer visual transformer ViT-B/16 [38] as the image encoder, and initialize it with weights pre-trained on ImageNet-1k from [31].”
    From this paper · §ALBEF Pre-training
Show 16 more
  • MS COCO2014 · cited 1×, 1 in Method
    “Following UNITER [2], we construct our pre-training data using two web datasets (Conceptual Captions [4], SBU Captions [5]) and two in-domain datasets (COCO [41] and Visual Genome [42]).”
    From this paper · §ALBEF Pre-training
  • Visual Genome2016 · cited 1×, 1 in Method
    “Following UNITER [2], we construct our pre-training data using two web datasets (Conceptual Captions [4], SBU Captions [5]) and two in-domain datasets (COCO [41] and Visual Genome [42]).”
    From this paper · §ALBEF Pre-training
  • Transformer2017 · cited 1×, 1 in Method
    “We use a 6-layer transformer [39] for both the text encoder and the multimodal encoder.”
    From this paper · §ALBEF Pre-training
  • AdamW2017 · cited 1×, 1 in Method
    “We use the AdamW [44] optimizer with a weight decay of 0.02.”
    From this paper · §ALBEF Pre-training
  • BERT2018 · cited 1×, 1 in Method
    “The text encoder is initialized using the first 6 layers of the BERTbase [40] model, and the multimodal encoder is initialized using the last 6 layers of the BERTbase.”
    From this paper · §ALBEF Pre-training
  • Conceptual 12M2021 · cited 1×, 1 in Method
    “To show that our method is scalable with larger-scale web data, we also include the much noisier Conceptual 12M dataset [43], increasing the total number of images to 14.1M 33 3 some urls provided by the web datasets hav…”
    From this paper · §ALBEF Pre-training
  • VILLA2020 · cited 6×
    “The first category focuses on modelling the interactions between image and text features with transformer-based multimodal encoders [10, 11, 12, 13, 1, 14, 15, 2, 3, 16, 8, 17, 18].”
    From this paper · §Related Work
  • Oscar2020 · cited 5×
    “We evaluate ALBEF on the Flickr30K [49] and COCO benchmarks, and fine-tune the pre-trained model using the training samples from each dataset.”
    From this paper · §Downstream V+L Tasks
  • ALIGN2021 · cited 4×
    “The first category focuses on modelling the interactions between image and text features with transformer-based multimodal encoders [10, 11, 12, 13, 1, 14, 15, 2, 3, 16, 8, 17, 18].”
    From this paper · §Related Work
  • CLIP2021 · cited 4×
    “The first category focuses on modelling the interactions between image and text features with transformer-based multimodal encoders [10, 11, 12, 13, 1, 14, 15, 2, 3, 16, 8, 17, 18].”
    From this paper · §Related Work
  • LXMERT2019 · cited 3×
    “The first category focuses on modelling the interactions between image and text features with transformer-based multimodal encoders [10, 11, 12, 13, 1, 14, 15, 2, 3, 16, 8, 17, 18].”
    From this paper · §Related Work
  • Knowledge Distillation2015 · cited 2×
    “Knowledge distillation [28] aims to improve a student model’s performance by distilling knowledge from a teacher model, usually through matching the student’s prediction with the teacher’s.”
    From this paper · §Related Work
  • NLVR22018 · cited 2×
    “NLVR2 [19], VQA [20]), but most of them require high-resolution input images and pre-trained object detectors.”
    From this paper · §Related Work
  • SimCLR2020 · cited 2×
    “The recent CLIP [6] and ALIGN [7] perform pre-training on massive noisy web data using a contrastive loss, one of the most effective loss for representation learning [24, 25, 26, 27].”
    From this paper · §Related Work
  • VinVL2021 · cited 2×
    “The first category focuses on modelling the interactions between image and text features with transformer-based multimodal encoders [10, 11, 12, 13, 1, 14, 15, 2, 3, 16, 8, 17, 18].”
    From this paper · §Related Work
  • ViLT2021 · cited 2×
    “A recent method [21] improves inference speed by removing the object detector, but results in lower performance.”
    From this paper · §Related Work

Led to

  • VLMo2021 · cited 9×, 2 in Method
    “Pre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”
    From VLMo · §Related Work
  • METER2021 · cited 6×, 4 in Method
    “Recent works kim2021vilt; xue2021probing; li2021align that tried to adopt vision transformers have not shown satisfactory performance and typically underperform state-of-the-art region-feature-based VLP models (e.g., Vin…”
    From METER · §Introduction
  • BLIP2022 · cited 14×, 5 in Method
    “Due to the prohibitive expense of acquiring human-annotated texts, most methods (Chen et al. 2020; Li et al. 2020; Li et al. 2021a; Wang et al. 2021; Radford et al. 2021) use image and alt-text pairs crawled from the web…”
    From BLIP · §Related Work
  • CoCa2022 · cited 4×, 2 in Method
    “Since unidirectional language models are trained with causal masking on complete sentences, the decoder can efficiently generate outputs for both contrastive and generative losses with a single forward propagation (compa…”
    From CoCa · §Approach
  • VL-BEiT2022 · cited 3×
    “Vision-language pretraining [37, 24, 35, 46, 29, 21, 16, 20, 42, 41, 40, 1, 45] aims to learn multimodal representations from large-scale image-text pairs.”
    From VL-BEiT · §Related Work
  • BEiT-32022 · cited 1×, 1 in Method
    “In contrast, previous vision-language models [36, 65, 26, 34, 53, 30, 62] usually employ multiple pretraining tasks, such as image-text contrast, image-text matching, and word-patch/region alignment.”
    From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • BLIP-22023 · cited 7×, 1 in Method
    “We adopt the hard negative mining strategy from Li et al. 2021; Li et al. 2022 to create informative negative pairs.”
    From BLIP-2 · §Method
Abstract

Large-scale vision and language representation learning has shown promising improvements on various vision-language tasks. Most existing methods employ a transformer-based multimodal encoder to jointly model visual tokens (region-based image features) and word tokens. Because the visual tokens and word tokens are unaligned, it is challenging for the multimodal encoder to learn image-text interactions. In this paper, we introduce a contrastive loss to ALign the image and text representations BEfore Fusing (ALBEF) them through cross-modal attention, which enables more grounded vision and language representation learning. Unlike most existing methods, our method does not require bounding box annotations nor high-resolution images. In order to improve learning from noisy web data, we propose momentum distillation, a self-training method which learns from pseudo-targets produced by a momentum model. We provide a theoretical analysis of ALBEF from a mutual information maximization perspective, showing that different training tasks can be interpreted as different ways to generate views for an image-text pair. ALBEF achieves state-of-the-art performance on multiple downstream vision-language tasks. On image-text retrieval, ALBEF outperforms methods that are pre-trained on orders of magnitude larger datasets. On VQA and NLVR$^2$, ALBEF achieves absolute improvements of 2.37% and 3.84% compared to the state-of-the-art, while enjoying faster inference speed. Code and pre-trained models are available at https://github.com/salesforce/ALBEF/.