Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
Large-scale vision and language representation learning has shown promising improvements on various vision-language tasks. Most existing methods employ a transformer-based multimodal encoder to jointly model visual tokens (region-based image features) and word tokens.
Also cited · not yet reviewed (20)
- ViT2020 · cited 2×, 2 in Method“We use a 12-layer visual transformer ViT-B/16 [38] as the image encoder, and initialize it with weights pre-trained on ImageNet-1k from [31].”From this paper · §ALBEF Pre-training
- UNITER2019 · cited 7×, 1 in Method“Following UNITER [2], we construct our pre-training data using two web datasets (Conceptual Captions [4], SBU Captions [5]) and two in-domain datasets (COCO [41] and Visual Genome [42]).”From this paper · §ALBEF Pre-training
- MoCo2019 · cited 3×, 1 in Method“Inspired by MoCo [24], we maintain two queues to store the most recent MM image-text representations from the momentum unimodal encoders.”From this paper · §ALBEF Pre-training
- DeiT2020 · cited 2×, 1 in Method“We use a 12-layer visual transformer ViT-B/16 [38] as the image encoder, and initialize it with weights pre-trained on ImageNet-1k from [31].”From this paper · §ALBEF Pre-training
Show 16 more
- MS COCO2014 · cited 1×, 1 in Method“Following UNITER [2], we construct our pre-training data using two web datasets (Conceptual Captions [4], SBU Captions [5]) and two in-domain datasets (COCO [41] and Visual Genome [42]).”From this paper · §ALBEF Pre-training
- Visual Genome2016 · cited 1×, 1 in Method“Following UNITER [2], we construct our pre-training data using two web datasets (Conceptual Captions [4], SBU Captions [5]) and two in-domain datasets (COCO [41] and Visual Genome [42]).”From this paper · §ALBEF Pre-training
- Transformer2017 · cited 1×, 1 in Method“We use a 6-layer transformer [39] for both the text encoder and the multimodal encoder.”From this paper · §ALBEF Pre-training
- AdamW2017 · cited 1×, 1 in Method“We use the AdamW [44] optimizer with a weight decay of 0.02.”From this paper · §ALBEF Pre-training
- BERT2018 · cited 1×, 1 in Method“The text encoder is initialized using the first 6 layers of the BERTbase [40] model, and the multimodal encoder is initialized using the last 6 layers of the BERTbase.”From this paper · §ALBEF Pre-training
- Conceptual 12M2021 · cited 1×, 1 in Method“To show that our method is scalable with larger-scale web data, we also include the much noisier Conceptual 12M dataset [43], increasing the total number of images to 14.1M 33 3 some urls provided by the web datasets hav…”From this paper · §ALBEF Pre-training
- VILLA2020 · cited 6דThe first category focuses on modelling the interactions between image and text features with transformer-based multimodal encoders [10, 11, 12, 13, 1, 14, 15, 2, 3, 16, 8, 17, 18].”From this paper · §Related Work
- Oscar2020 · cited 5דWe evaluate ALBEF on the Flickr30K [49] and COCO benchmarks, and fine-tune the pre-trained model using the training samples from each dataset.”From this paper · §Downstream V+L Tasks
- ALIGN2021 · cited 4דThe first category focuses on modelling the interactions between image and text features with transformer-based multimodal encoders [10, 11, 12, 13, 1, 14, 15, 2, 3, 16, 8, 17, 18].”From this paper · §Related Work
- CLIP2021 · cited 4דThe first category focuses on modelling the interactions between image and text features with transformer-based multimodal encoders [10, 11, 12, 13, 1, 14, 15, 2, 3, 16, 8, 17, 18].”From this paper · §Related Work
- LXMERT2019 · cited 3דThe first category focuses on modelling the interactions between image and text features with transformer-based multimodal encoders [10, 11, 12, 13, 1, 14, 15, 2, 3, 16, 8, 17, 18].”From this paper · §Related Work
- Knowledge Distillation2015 · cited 2דKnowledge distillation [28] aims to improve a student model’s performance by distilling knowledge from a teacher model, usually through matching the student’s prediction with the teacher’s.”From this paper · §Related Work
- NLVR22018 · cited 2דNLVR2 [19], VQA [20]), but most of them require high-resolution input images and pre-trained object detectors.”From this paper · §Related Work
- SimCLR2020 · cited 2דThe recent CLIP [6] and ALIGN [7] perform pre-training on massive noisy web data using a contrastive loss, one of the most effective loss for representation learning [24, 25, 26, 27].”From this paper · §Related Work
- VinVL2021 · cited 2דThe first category focuses on modelling the interactions between image and text features with transformer-based multimodal encoders [10, 11, 12, 13, 1, 14, 15, 2, 3, 16, 8, 17, 18].”From this paper · §Related Work
- ViLT2021 · cited 2דA recent method [21] improves inference speed by removing the object detector, but results in lower performance.”From this paper · §Related Work
Led to
- VLMo2021 · cited 9×, 2 in Method“Pre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”From VLMo · §Related Work
- METER2021 · cited 6×, 4 in Method“Recent works kim2021vilt; xue2021probing; li2021align that tried to adopt vision transformers have not shown satisfactory performance and typically underperform state-of-the-art region-feature-based VLP models (e.g., Vin…”From METER · §Introduction
- BLIP2022 · cited 14×, 5 in Method“Due to the prohibitive expense of acquiring human-annotated texts, most methods (Chen et al. 2020; Li et al. 2020; Li et al. 2021a; Wang et al. 2021; Radford et al. 2021) use image and alt-text pairs crawled from the web…”From BLIP · §Related Work
- CoCa2022 · cited 4×, 2 in Method“Since unidirectional language models are trained with causal masking on complete sentences, the decoder can efficiently generate outputs for both contrastive and generative losses with a single forward propagation (compa…”From CoCa · §Approach
- VL-BEiT2022 · cited 3דVision-language pretraining [37, 24, 35, 46, 29, 21, 16, 20, 42, 41, 40, 1, 45] aims to learn multimodal representations from large-scale image-text pairs.”From VL-BEiT · §Related Work
- BEiT-32022 · cited 1×, 1 in Method“In contrast, previous vision-language models [36, 65, 26, 34, 53, 30, 62] usually employ multiple pretraining tasks, such as image-text contrast, image-text matching, and word-patch/region alignment.”From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
- BLIP-22023 · cited 7×, 1 in Method“We adopt the hard negative mining strategy from Li et al. 2021; Li et al. 2022 to create informative negative pairs.”From BLIP-2 · §Method
Abstract
Large-scale vision and language representation learning has shown promising improvements on various vision-language tasks. Most existing methods employ a transformer-based multimodal encoder to jointly model visual tokens (region-based image features) and word tokens. Because the visual tokens and word tokens are unaligned, it is challenging for the multimodal encoder to learn image-text interactions. In this paper, we introduce a contrastive loss to ALign the image and text representations BEfore Fusing (ALBEF) them through cross-modal attention, which enables more grounded vision and language representation learning. Unlike most existing methods, our method does not require bounding box annotations nor high-resolution images. In order to improve learning from noisy web data, we propose momentum distillation, a self-training method which learns from pseudo-targets produced by a momentum model. We provide a theoretical analysis of ALBEF from a mutual information maximization perspective, showing that different training tasks can be interpreted as different ways to generate views for an image-text pair. ALBEF achieves state-of-the-art performance on multiple downstream vision-language tasks. On image-text retrieval, ALBEF outperforms methods that are pre-trained on orders of magnitude larger datasets. On VQA and NLVR$^2$, ALBEF achieves absolute improvements of 2.37% and 3.84% compared to the state-of-the-art, while enjoying faster inference speed. Code and pre-trained models are available at https://github.com/salesforce/ALBEF/.