Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training
We propose Unicoder-VL, a universal encoder that aims to learn joint representations of vision and language in a pre-training manner. Borrow ideas from cross-lingual pre-trained models, such as XLM and Unicoder, both visual and linguistic contents are fed into a multi-layer Transformer for the cross-modal pre-training, where three pre-trained tasks are employed, including Masked Language Modeling (MLM), Masked Object Classification (MOC) and Visual-linguistic Matching (VLM).
Also cited · not yet reviewed (5)
- BERT2018 · cited 3×, 1 in Method“BERT [\citeauthoryearDevlin et al.2018] is a pre-trained model based on multi-layer Transformer [\citeauthoryearVaswani et al.2017].”From this paper · §Approach
- GNMT2016 · cited 1×, 1 in Method“TT is the length of the WordPiece [\citeauthoryearWu et al.2016] linguistic input.”From this paper · §Approach
- ViLBERT2019 · cited 4דHowever, for image RoI based methods like SCAN[\citeauthoryearLee et al.2018], Unicoder-VL and ViLBERT [\citeauthoryearLu et al.2019], the backbone of Faster-RCNN is still not fine-tuned with the whole model during cross…”From this paper · §Results and Analysis
- UNITER2019 · cited 4דConcurrent to our work, several recent released works, such as ViLBERT [\citeauthoryearLu et al.2019], VisualBERT [\citeauthoryearLi et al.2019], VL-BERT [\citeauthoryearSu et al.2019] and UNITER [\citeauthoryearChen et…”From this paper · §Related Work
Show 1 more
- RoBERTa2019 · cited 2דLatest pre-trained NLP models are based on multi-layer Transformer, such as GPT [\citeauthoryearRadford et al.2018], BERT [\citeauthoryearDevlin et al.2018], XLNet (Yang et al., 2019) and RoBERTa[\citeauthoryearLiu et al…”From this paper · §Related Work
Led to
- 12-in-12019 · cited 6×, 2 in Method“The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”From 12-in-1 · §Introduction
- Oscar2020 · cited 3דFollowing [19], we report the top-KK retrieval results on both the 11K and 55K COCO test sets.”From Oscar · §Adapting to V+L Tasks
- VL-BERT meta-analysis2020 · cited 2דThe majority of V&L BERTs follow the single-stream paradigm (Su et al. 2020; Li et al. 2019; Chen et al. 2020; Li et al. 2020a; Zhou et al. 2020; Lin et al. 2020; Li et al. 2020b).”From VL-BERT meta-analysis · §Vision-and-Language BERTs
- VinVL2021 · cited 2דFollowing [19], we report the top-KK retrieval results on both the 11K and 55K COCO test sets.”From VinVL · §Adapting to VL Tasks
- ViLT2021 · cited 2×, 1 in Method“Backbone: ResNet-101 (Lu et al. 2019; Tan & Bansal 2019; Su et al. 2019) and ResNext-152 (Li et al. 2019; Li et al. 2020a; Zhang et al. 2021) are two commonly used backbones.”From ViLT · §Background
- Conceptual 12M2021 · cited 6×, 1 in Method“The large scale nature and the high degree of textual and visual diversity make this dataset particularly suited to V+L pre-training [55, 21, 74, 88, 48, 56].”From Conceptual 12M · §Vision-and-Language Pre-Training Data
Abstract
We propose Unicoder-VL, a universal encoder that aims to learn joint representations of vision and language in a pre-training manner. Borrow ideas from cross-lingual pre-trained models, such as XLM and Unicoder, both visual and linguistic contents are fed into a multi-layer Transformer for the cross-modal pre-training, where three pre-trained tasks are employed, including Masked Language Modeling (MLM), Masked Object Classification (MOC) and Visual-linguistic Matching (VLM). The first two tasks learn context-aware representations for input tokens based on linguistic and visual contents jointly. The last task tries to predict whether an image and a text describe each other. After pretraining on large-scale image-caption pairs, we transfer Unicoder-VL to caption-based image-text retrieval and visual commonsense reasoning, with just one additional output layer. We achieve state-of-the-art or comparable results on both two tasks and show the powerful ability of the cross-modal pre-training.