Unified Vision-Language Pre-Training for Image Captioning and VQA
This paper presents a unified Vision-Language Pre-training (VLP) model. , visual question answering) tasks, and (2) it uses a shared multi-layer transformer network for both encoding and decoding, which differs from many existing methods where the encoder and decoder are implemented using separate models.
Also cited · not yet reviewed (6)
- BERT2018 · cited 7×, 1 in Method“Inspired by the recent success of pre-trained language models such as BERT [\citeauthoryearDevlin et al.2018] and GPT [\citeauthoryearRadford et al.2018, \citeauthoryearRadford et al.2019], there is a growing interest in…”From this paper · §Introduction
- LXMERT2019 · cited 6×, 1 in Method“Inspired by the recent success of pre-trained language models such as BERT [\citeauthoryearDevlin et al.2018] and GPT [\citeauthoryearRadford et al.2018, \citeauthoryearRadford et al.2019], there is a growing interest in…”From this paper · §Introduction
- UniLM2019 · cited 4×, 1 in Method“We follow the same scheme and consider two specific objectives: the bidirectional objective (bidirectional) as in BERT and the sequence to sequence objective (seq2seq), inspired by [\citeauthoryearDong et al.2019].”From this paper · §Vision-Language Pre-training
- ViLBERT2019 · cited 4×, 1 in Method“Inspired by the recent success of pre-trained language models such as BERT [\citeauthoryearDevlin et al.2018] and GPT [\citeauthoryearRadford et al.2018, \citeauthoryearRadford et al.2019], there is a growing interest in…”From this paper · §Introduction
Show 2 more
- RoBERTa2019 · cited 1×, 1 in Method“This coincidentally agrees with a concurrent work of RoBERTa [\citeauthoryearLiu et al.2019b].”From this paper · §Vision-Language Pre-training
- COCO Captions2015 · cited 2דWe validate VLP in our experiments on both the image captioning and VQA tasks using three challenging benchmarks: COCO Captions [\citeauthoryearChen et al.2015], Flickr30k Captions [\citeauthoryearYoung et al.2014], and…”From this paper · §Introduction
Led to
- 12-in-12019 · cited 3×, 1 in Method“The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”From 12-in-1 · §Introduction
- VL-BERT meta-analysis2020 · cited 2דThe majority of V&L BERTs follow the single-stream paradigm (Su et al. 2020; Li et al. 2019; Chen et al. 2020; Li et al. 2020a; Zhou et al. 2020; Lin et al. 2020; Li et al. 2020b).”From VL-BERT meta-analysis · §Vision-and-Language BERTs
- VinVL2021 · cited 2דSimilar to VLP [21, 45], the self-attention mask is constrained such that a caption token can only attend to the tokens before its position to simulate a uni-directional generation process.”From VinVL · §Adapting to VL Tasks
- VL-T52021 · cited 2דIn this section, we compare our generative architectures VL-T5 and VL-BART on a diverse set of 7 downstream tasks (details in Appendix) with existing vision-and-language pretrained transformers (Tan & Bansal 2019; Lu et…”From VL-T5 · §Downstream Tasks and Results
- Conceptual 12M2021 · cited 7×, 2 in Method“The large scale nature and the high degree of textual and visual diversity make this dataset particularly suited to V+L pre-training [55, 21, 74, 88, 48, 56].”From Conceptual 12M · §Vision-and-Language Pre-Training Data
- ViTCAP2021 · cited 2דRecent studies have witnessed its great development which are primarily reflected in the aspects of more advanced cross-modal fusion architectures you2016image; vinyals2015show; xu2015show; rennie2017self; yang2019auto;…”From ViTCAP · §Introduction
- BLIP2022 · cited 2דThere have been many attempts to unify various vision and language tasks into a single framework (Zhou et al. 2020; Cho et al. 2021; Wang et al. 2021).”From BLIP · §Related Work
- Simple end-to-end captioning2022 · cited 1×, 1 in Method“(2) Large-scale cross-modal pre-trained models: we include the results of 7 cross-modal pre-trained models, including UVLP (Zhou et al. 2019b), OSCAR (Li et al. 2020), XGPT (Xia et al. 2021), MiniVLM (Wang et al. 2020),…”From Simple end-to-end captioning · §. Experiment Setup
Abstract
This paper presents a unified Vision-Language Pre-training (VLP) model. The model is unified in that (1) it can be fine-tuned for either vision-language generation (e.g., image captioning) or understanding (e.g., visual question answering) tasks, and (2) it uses a shared multi-layer transformer network for both encoding and decoding, which differs from many existing methods where the encoder and decoder are implemented using separate models. The unified VLP model is pre-trained on a large amount of image-text pairs using the unsupervised learning objectives of two tasks: bidirectional and sequence-to-sequence (seq2seq) masked vision-language prediction. The two tasks differ solely in what context the prediction conditions on. This is controlled by utilizing specific self-attention masks for the shared transformer network. To the best of our knowledge, VLP is the first reported model that achieves state-of-the-art results on both vision-language generation and understanding tasks, as disparate as image captioning and visual question answering, across three challenging benchmark datasets: COCO Captions, Flickr30k Captions, and VQA 2.0. The code and the pre-trained models are available at https://github.com/LuoweiZhou/VLP.