UNITER: UNiversal Image-TExt Representation Learning
Joint image-text embedding is the bedrock for most Vision-and-Language (V+L) tasks, where multimodality inputs are simultaneously processed for joint visual and textual understanding. In this paper, we introduce UNITER, a UNiversal Image-TExt Representation, learned through large-scale pre-training over four image-text datasets (COCO, Visual Genome, Conceptual Captions, and SBU Captions), which can power heterogeneous downstream V+L tasks with joint multimodal embeddings.
Also cited · not yet reviewed (1)
- VisualBERT2019 · cited 3דOn the other hand, B2T2 [1], VisualBERT [25], Unicoder-VL [24] and VL-BERT [42] proposed the single-stream architecture, where a single Transformer is applied to both images and text.”From this paper · §Related Work
Led to
- Unicoder-VL2019 · cited 4דConcurrent to our work, several recent released works, such as ViLBERT [\citeauthoryearLu et al.2019], VisualBERT [\citeauthoryearLi et al.2019], VL-BERT [\citeauthoryearSu et al.2019] and UNITER [\citeauthoryearChen et…”From Unicoder-VL · §Related Work
- 12-in-12019 · cited 4×, 1 in Method“Recent work has used online hard-negative mining chen2019uniter; li2019unicoder but this is costly to compute.”From 12-in-1 · §Approach
- Oscar2020 · cited 5דThe existing methods [37, 38, 22, 5, 46, 35, 19, 10] employ BERT-like objectives [6] to learn cross-modal representations from a concatenated-sequence of visual region features and language token embeddings.”From Oscar · §Related Work
- VILLA2020 · cited 4דInspired by the success of BERT [13] on natural language understanding, there has been a surging research interest in developing multimodal pre-training methods for vision-and-language representation learning (e.g., ViLB…”From VILLA · §Introduction
- VirTex2020 · cited 2דInspired by the success of BERT [64] in NLP, several recent methods use Transformers [29] to learn transferable joint representations of images and text [65, 66, 67, 68, 69, 70, 71, 72].”From VirTex · §Related Work
- VL-BERT meta-analysis2020 · cited 7דThe majority of V&L BERTs follow the single-stream paradigm (Su et al. 2020; Li et al. 2019; Chen et al. 2020; Li et al. 2020a; Zhou et al. 2020; Lin et al. 2020; Li et al. 2020b).”From VL-BERT meta-analysis · §Vision-and-Language BERTs
- VL-T52021 · cited 12×, 4 in Method“As shown in Fig.3 (a), existing methods (Tan & Bansal 2019; Lu et al. 2019; Chen et al. 2020) typically introduce a multi-layer perceptron (MLP) multi-label classifier head on top of h[CLS]xh^{x}_{\texttt{[CLS]}}, which…”From VL-T5 · §Model
- ViLT2021 · cited 5×, 1 in Method“Plus, inspired by the word region alignment objective in Chen et al. 2019, we design word patch alignment (WPA) that computes the alignment score between two subsets of zDz^{D}: zD|tz^{D}|_{t} (textual subset) and zD|vz^…”From ViLT · §Vision-and-Language Transformer
- ALIGN2021 · cited 3דRecently more advanced models emerge with cross-modal attention layers (Liu et al. 2019a; Lu et al. 2019; Chen et al. 2020c; Huang et al. 2020b) and show superior performance in image-text matching tasks.”From ALIGN · §Related Work
- Conceptual 12M2021 · cited 7×, 1 in Method“The large scale nature and the high degree of textual and visual diversity make this dataset particularly suited to V+L pre-training [55, 21, 74, 88, 48, 56].”From Conceptual 12M · §Vision-and-Language Pre-Training Data
- ALBEF2021 · cited 7×, 1 in Method“Following UNITER [2], we construct our pre-training data using two web datasets (Conceptual Captions [4], SBU Captions [5]) and two in-domain datasets (COCO [41] and Visual Genome [42]).”From ALBEF · §ALBEF Pre-training
- SimVLM2021 · cited 5דWhile a variety of approaches have been proposed, a large portion of them require object detection for image region feature regression or tagging as part of the pre-training objectives (Tan & Bansal 2019; Su et al. 2020;…”From SimVLM · §Related Work
- VLMo2021 · cited 6דPre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”From VLMo · §Related Work
- METER2021 · cited 11×, 9 in Method“Vision-and-language pre-training (VLP) has now become the de facto practice to tackle these tasks tan-bansal-2019-lxmert; li2019visualbert; lu2019vilbert; su2019vl; chen2020uniter; li2020oscar.”From METER · §Introduction
- Florence2021 · cited 1×, 1 in Method“Thus, the object detector has been a de facto tool for image feature extraction, followed by a fusion network for prediction in many works (Anderson et al. 2018; Li et al. 2020; Zhang et al. 2021b; Wang et al. 2020; Fang…”From Florence · §Approach
- BLIP2022 · cited 3×, 1 in Method“Due to the prohibitive expense of acquiring human-annotated texts, most methods (Chen et al. 2020; Li et al. 2020; Li et al. 2021a; Wang et al. 2021; Radford et al. 2021) use image and alt-text pairs crawled from the web…”From BLIP · §Related Work
Abstract
Joint image-text embedding is the bedrock for most Vision-and-Language (V+L) tasks, where multimodality inputs are simultaneously processed for joint visual and textual understanding. In this paper, we introduce UNITER, a UNiversal Image-TExt Representation, learned through large-scale pre-training over four image-text datasets (COCO, Visual Genome, Conceptual Captions, and SBU Captions), which can power heterogeneous downstream V+L tasks with joint multimodal embeddings. We design four pre-training tasks: Masked Language Modeling (MLM), Masked Region Modeling (MRM, with three variants), Image-Text Matching (ITM), and Word-Region Alignment (WRA). Different from previous work that applies joint random masking to both modalities, we use conditional masking on pre-training tasks (i.e., masked language/region modeling is conditioned on full observation of image/text). In addition to ITM for global image-text alignment, we also propose WRA via the use of Optimal Transport (OT) to explicitly encourage fine-grained alignment between words and image regions during pre-training. Comprehensive analysis shows that both conditional masking and OT-based WRA contribute to better pre-training. We also conduct a thorough ablation study to find an optimal combination of pre-training tasks. Extensive experiments show that UNITER achieves new state of the art across six V+L tasks (over nine datasets), including Visual Question Answering, Image-Text Retrieval, Referring Expression Comprehension, Visual Commonsense Reasoning, Visual Entailment, and NLVR$^2$. Code is available at https://github.com/ChenRocks/UNITER.