Paper Lineage
Esc
MethodSep 2019arXiv 1909.11740cs.CV

UNITER: UNiversal Image-TExt Representation Learning

Yen-Chun Chen, Linjie Li, Licheng Yu and 5 others

Joint image-text embedding is the bedrock for most Vision-and-Language (V+L) tasks, where multimodality inputs are simultaneously processed for joint visual and textual understanding. In this paper, we introduce UNITER, a UNiversal Image-TExt Representation, learned through large-scale pre-training over four image-text datasets (COCO, Visual Genome, Conceptual Captions, and SBU Captions), which can power heterogeneous downstream V+L tasks with joint multimodal embeddings.

From the abstract

Built on

1 paper · 0 verifiedSee as graph

Also cited · not yet reviewed (1)

  • VisualBERT2019 · cited 3×
    “On the other hand, B2T2 [1], VisualBERT [25], Unicoder-VL [24] and VL-BERT [42] proposed the single-stream architecture, where a single Transformer is applied to both images and text.”
    From this paper · §Related Work

Led to

  • Unicoder-VL2019 · cited 4×
    “Concurrent to our work, several recent released works, such as ViLBERT [\citeauthoryearLu et al.2019], VisualBERT [\citeauthoryearLi et al.2019], VL-BERT [\citeauthoryearSu et al.2019] and UNITER [\citeauthoryearChen et…”
    From Unicoder-VL · §Related Work
  • 12-in-12019 · cited 4×, 1 in Method
    “Recent work has used online hard-negative mining chen2019uniter; li2019unicoder but this is costly to compute.”
    From 12-in-1 · §Approach
  • Oscar2020 · cited 5×
    “The existing methods [37, 38, 22, 5, 46, 35, 19, 10] employ BERT-like objectives [6] to learn cross-modal representations from a concatenated-sequence of visual region features and language token embeddings.”
    From Oscar · §Related Work
  • VILLA2020 · cited 4×
    “Inspired by the success of BERT [13] on natural language understanding, there has been a surging research interest in developing multimodal pre-training methods for vision-and-language representation learning (e.g., ViLB…”
    From VILLA · §Introduction
  • VirTex2020 · cited 2×
    “Inspired by the success of BERT [64] in NLP, several recent methods use Transformers [29] to learn transferable joint representations of images and text [65, 66, 67, 68, 69, 70, 71, 72].”
    From VirTex · §Related Work
  • VL-BERT meta-analysis2020 · cited 7×
    “The majority of V&L BERTs follow the single-stream paradigm (Su et al. 2020; Li et al. 2019; Chen et al. 2020; Li et al. 2020a; Zhou et al. 2020; Lin et al. 2020; Li et al. 2020b).”
    From VL-BERT meta-analysis · §Vision-and-Language BERTs
  • VL-T52021 · cited 12×, 4 in Method
    “As shown in Fig.3 (a), existing methods (Tan & Bansal 2019; Lu et al. 2019; Chen et al. 2020) typically introduce a multi-layer perceptron (MLP) multi-label classifier head on top of h[CLS]xh^{x}_{\texttt{[CLS]}}, which…”
    From VL-T5 · §Model
  • ViLT2021 · cited 5×, 1 in Method
    “Plus, inspired by the word region alignment objective in Chen et al. 2019, we design word patch alignment (WPA) that computes the alignment score between two subsets of zDz^{D}: zD|tz^{D}|_{t} (textual subset) and zD|vz^…”
    From ViLT · §Vision-and-Language Transformer
  • ALIGN2021 · cited 3×
    “Recently more advanced models emerge with cross-modal attention layers (Liu et al. 2019a; Lu et al. 2019; Chen et al. 2020c; Huang et al. 2020b) and show superior performance in image-text matching tasks.”
    From ALIGN · §Related Work
  • Conceptual 12M2021 · cited 7×, 1 in Method
    “The large scale nature and the high degree of textual and visual diversity make this dataset particularly suited to V+L pre-training [55, 21, 74, 88, 48, 56].”
    From Conceptual 12M · §Vision-and-Language Pre-Training Data
  • ALBEF2021 · cited 7×, 1 in Method
    “Following UNITER [2], we construct our pre-training data using two web datasets (Conceptual Captions [4], SBU Captions [5]) and two in-domain datasets (COCO [41] and Visual Genome [42]).”
    From ALBEF · §ALBEF Pre-training
  • SimVLM2021 · cited 5×
    “While a variety of approaches have been proposed, a large portion of them require object detection for image region feature regression or tagging as part of the pre-training objectives (Tan & Bansal 2019; Su et al. 2020;…”
    From SimVLM · §Related Work
  • VLMo2021 · cited 6×
    “Pre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”
    From VLMo · §Related Work
  • METER2021 · cited 11×, 9 in Method
    “Vision-and-language pre-training (VLP) has now become the de facto practice to tackle these tasks tan-bansal-2019-lxmert; li2019visualbert; lu2019vilbert; su2019vl; chen2020uniter; li2020oscar.”
    From METER · §Introduction
  • Florence2021 · cited 1×, 1 in Method
    “Thus, the object detector has been a de facto tool for image feature extraction, followed by a fusion network for prediction in many works (Anderson et al. 2018; Li et al. 2020; Zhang et al. 2021b; Wang et al. 2020; Fang…”
    From Florence · §Approach
  • BLIP2022 · cited 3×, 1 in Method
    “Due to the prohibitive expense of acquiring human-annotated texts, most methods (Chen et al. 2020; Li et al. 2020; Li et al. 2021a; Wang et al. 2021; Radford et al. 2021) use image and alt-text pairs crawled from the web…”
    From BLIP · §Related Work
Abstract

Joint image-text embedding is the bedrock for most Vision-and-Language (V+L) tasks, where multimodality inputs are simultaneously processed for joint visual and textual understanding. In this paper, we introduce UNITER, a UNiversal Image-TExt Representation, learned through large-scale pre-training over four image-text datasets (COCO, Visual Genome, Conceptual Captions, and SBU Captions), which can power heterogeneous downstream V+L tasks with joint multimodal embeddings. We design four pre-training tasks: Masked Language Modeling (MLM), Masked Region Modeling (MRM, with three variants), Image-Text Matching (ITM), and Word-Region Alignment (WRA). Different from previous work that applies joint random masking to both modalities, we use conditional masking on pre-training tasks (i.e., masked language/region modeling is conditioned on full observation of image/text). In addition to ITM for global image-text alignment, we also propose WRA via the use of Optimal Transport (OT) to explicitly encourage fine-grained alignment between words and image regions during pre-training. Comprehensive analysis shows that both conditional masking and OT-based WRA contribute to better pre-training. We also conduct a thorough ablation study to find an optimal combination of pre-training tasks. Extensive experiments show that UNITER achieves new state of the art across six V+L tasks (over nine datasets), including Visual Question Answering, Image-Text Retrieval, Referring Expression Comprehension, Visual Commonsense Reasoning, Visual Entailment, and NLVR$^2$. Code is available at https://github.com/ChenRocks/UNITER.