Paper Lineage
Esc
MethodSep 2019arXiv 1909.11059cs.CV

Unified Vision-Language Pre-Training for Image Captioning and VQA

Luowei Zhou, Hamid Palangi, Lei Zhang and 3 others

This paper presents a unified Vision-Language Pre-training (VLP) model. , visual question answering) tasks, and (2) it uses a shared multi-layer transformer network for both encoding and decoding, which differs from many existing methods where the encoder and decoder are implemented using separate models.

From the abstract

Built on

6 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (6)

  • BERT2018 · cited 7×, 1 in Method
    “Inspired by the recent success of pre-trained language models such as BERT [\citeauthoryearDevlin et al.2018] and GPT [\citeauthoryearRadford et al.2018, \citeauthoryearRadford et al.2019], there is a growing interest in…”
    From this paper · §Introduction
  • LXMERT2019 · cited 6×, 1 in Method
    “Inspired by the recent success of pre-trained language models such as BERT [\citeauthoryearDevlin et al.2018] and GPT [\citeauthoryearRadford et al.2018, \citeauthoryearRadford et al.2019], there is a growing interest in…”
    From this paper · §Introduction
  • UniLM2019 · cited 4×, 1 in Method
    “We follow the same scheme and consider two specific objectives: the bidirectional objective (bidirectional) as in BERT and the sequence to sequence objective (seq2seq), inspired by [\citeauthoryearDong et al.2019].”
    From this paper · §Vision-Language Pre-training
  • ViLBERT2019 · cited 4×, 1 in Method
    “Inspired by the recent success of pre-trained language models such as BERT [\citeauthoryearDevlin et al.2018] and GPT [\citeauthoryearRadford et al.2018, \citeauthoryearRadford et al.2019], there is a growing interest in…”
    From this paper · §Introduction
Show 2 more
  • RoBERTa2019 · cited 1×, 1 in Method
    “This coincidentally agrees with a concurrent work of RoBERTa [\citeauthoryearLiu et al.2019b].”
    From this paper · §Vision-Language Pre-training
  • COCO Captions2015 · cited 2×
    “We validate VLP in our experiments on both the image captioning and VQA tasks using three challenging benchmarks: COCO Captions [\citeauthoryearChen et al.2015], Flickr30k Captions [\citeauthoryearYoung et al.2014], and…”
    From this paper · §Introduction

Led to

  • 12-in-12019 · cited 3×, 1 in Method
    “The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”
    From 12-in-1 · §Introduction
  • VL-BERT meta-analysis2020 · cited 2×
    “The majority of V&L BERTs follow the single-stream paradigm (Su et al. 2020; Li et al. 2019; Chen et al. 2020; Li et al. 2020a; Zhou et al. 2020; Lin et al. 2020; Li et al. 2020b).”
    From VL-BERT meta-analysis · §Vision-and-Language BERTs
  • VinVL2021 · cited 2×
    “Similar to VLP [21, 45], the self-attention mask is constrained such that a caption token can only attend to the tokens before its position to simulate a uni-directional generation process.”
    From VinVL · §Adapting to VL Tasks
  • VL-T52021 · cited 2×
    “In this section, we compare our generative architectures VL-T5 and VL-BART on a diverse set of 7 downstream tasks (details in Appendix) with existing vision-and-language pretrained transformers (Tan & Bansal 2019; Lu et…”
    From VL-T5 · §Downstream Tasks and Results
  • Conceptual 12M2021 · cited 7×, 2 in Method
    “The large scale nature and the high degree of textual and visual diversity make this dataset particularly suited to V+L pre-training [55, 21, 74, 88, 48, 56].”
    From Conceptual 12M · §Vision-and-Language Pre-Training Data
  • ViTCAP2021 · cited 2×
    “Recent studies have witnessed its great development which are primarily reflected in the aspects of more advanced cross-modal fusion architectures you2016image; vinyals2015show; xu2015show; rennie2017self; yang2019auto;…”
    From ViTCAP · §Introduction
  • BLIP2022 · cited 2×
    “There have been many attempts to unify various vision and language tasks into a single framework (Zhou et al. 2020; Cho et al. 2021; Wang et al. 2021).”
    From BLIP · §Related Work
  • Simple end-to-end captioning2022 · cited 1×, 1 in Method
    “(2) Large-scale cross-modal pre-trained models: we include the results of 7 cross-modal pre-trained models, including UVLP (Zhou et al. 2019b), OSCAR (Li et al. 2020), XGPT (Xia et al. 2021), MiniVLM (Wang et al. 2020),…”
    From Simple end-to-end captioning · §. Experiment Setup
Abstract

This paper presents a unified Vision-Language Pre-training (VLP) model. The model is unified in that (1) it can be fine-tuned for either vision-language generation (e.g., image captioning) or understanding (e.g., visual question answering) tasks, and (2) it uses a shared multi-layer transformer network for both encoding and decoding, which differs from many existing methods where the encoder and decoder are implemented using separate models. The unified VLP model is pre-trained on a large amount of image-text pairs using the unsupervised learning objectives of two tasks: bidirectional and sequence-to-sequence (seq2seq) masked vision-language prediction. The two tasks differ solely in what context the prediction conditions on. This is controlled by utilizing specific self-attention masks for the shared transformer network. To the best of our knowledge, VLP is the first reported model that achieves state-of-the-art results on both vision-language generation and understanding tasks, as disparate as image captioning and visual question answering, across three challenging benchmark datasets: COCO Captions, Flickr30k Captions, and VQA 2.0. The code and the pre-trained models are available at https://github.com/LuoweiZhou/VLP.