Paper Lineage
Esc
MethodAug 2019arXiv 1908.06066cs.CV

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training

Gen Li, Nan Duan, Yuejian Fang and 3 others

We propose Unicoder-VL, a universal encoder that aims to learn joint representations of vision and language in a pre-training manner. Borrow ideas from cross-lingual pre-trained models, such as XLM and Unicoder, both visual and linguistic contents are fed into a multi-layer Transformer for the cross-modal pre-training, where three pre-trained tasks are employed, including Masked Language Modeling (MLM), Masked Object Classification (MOC) and Visual-linguistic Matching (VLM).

From the abstract

Built on

5 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (5)

  • BERT2018 · cited 3×, 1 in Method
    “BERT [\citeauthoryearDevlin et al.2018] is a pre-trained model based on multi-layer Transformer [\citeauthoryearVaswani et al.2017].”
    From this paper · §Approach
  • GNMT2016 · cited 1×, 1 in Method
    “TT is the length of the WordPiece [\citeauthoryearWu et al.2016] linguistic input.”
    From this paper · §Approach
  • ViLBERT2019 · cited 4×
    “However, for image RoI based methods like SCAN[\citeauthoryearLee et al.2018], Unicoder-VL and ViLBERT [\citeauthoryearLu et al.2019], the backbone of Faster-RCNN is still not fine-tuned with the whole model during cross…”
    From this paper · §Results and Analysis
  • UNITER2019 · cited 4×
    “Concurrent to our work, several recent released works, such as ViLBERT [\citeauthoryearLu et al.2019], VisualBERT [\citeauthoryearLi et al.2019], VL-BERT [\citeauthoryearSu et al.2019] and UNITER [\citeauthoryearChen et…”
    From this paper · §Related Work
Show 1 more
  • RoBERTa2019 · cited 2×
    “Latest pre-trained NLP models are based on multi-layer Transformer, such as GPT [\citeauthoryearRadford et al.2018], BERT [\citeauthoryearDevlin et al.2018], XLNet (Yang et al., 2019) and RoBERTa[\citeauthoryearLiu et al…”
    From this paper · §Related Work

Led to

  • 12-in-12019 · cited 6×, 2 in Method
    “The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”
    From 12-in-1 · §Introduction
  • Oscar2020 · cited 3×
    “Following [19], we report the top-KK retrieval results on both the 11K and 55K COCO test sets.”
    From Oscar · §Adapting to V+L Tasks
  • VL-BERT meta-analysis2020 · cited 2×
    “The majority of V&L BERTs follow the single-stream paradigm (Su et al. 2020; Li et al. 2019; Chen et al. 2020; Li et al. 2020a; Zhou et al. 2020; Lin et al. 2020; Li et al. 2020b).”
    From VL-BERT meta-analysis · §Vision-and-Language BERTs
  • VinVL2021 · cited 2×
    “Following [19], we report the top-KK retrieval results on both the 11K and 55K COCO test sets.”
    From VinVL · §Adapting to VL Tasks
  • ViLT2021 · cited 2×, 1 in Method
    “Backbone: ResNet-101 (Lu et al. 2019; Tan & Bansal 2019; Su et al. 2019) and ResNext-152 (Li et al. 2019; Li et al. 2020a; Zhang et al. 2021) are two commonly used backbones.”
    From ViLT · §Background
  • Conceptual 12M2021 · cited 6×, 1 in Method
    “The large scale nature and the high degree of textual and visual diversity make this dataset particularly suited to V+L pre-training [55, 21, 74, 88, 48, 56].”
    From Conceptual 12M · §Vision-and-Language Pre-Training Data
Abstract

We propose Unicoder-VL, a universal encoder that aims to learn joint representations of vision and language in a pre-training manner. Borrow ideas from cross-lingual pre-trained models, such as XLM and Unicoder, both visual and linguistic contents are fed into a multi-layer Transformer for the cross-modal pre-training, where three pre-trained tasks are employed, including Masked Language Modeling (MLM), Masked Object Classification (MOC) and Visual-linguistic Matching (VLM). The first two tasks learn context-aware representations for input tokens based on linguistic and visual contents jointly. The last task tries to predict whether an image and a text describe each other. After pretraining on large-scale image-caption pairs, we transfer Unicoder-VL to caption-based image-text retrieval and visual commonsense reasoning, with just one additional output layer. We achieve state-of-the-art or comparable results on both two tasks and show the powerful ability of the cross-modal pre-training.