Paper Lineage
Esc
MethodFeb 2021arXiv 2102.02779cs.CL

Unifying Vision-and-Language Tasks via Text Generation

Jaemin Cho, Jie Lei, Hao Tan, Mohit Bansal

Existing methods for vision-and-language learning typically require designing task-specific architectures and objectives for each task. For example, a multi-label answer classifier for visual question answering, a region scorer for referring expression comprehension, and a language decoder for image captioning, etc.

From the abstract

Built on

18 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (18)

  • UNITER2019 · cited 12×, 4 in Method
    “As shown in Fig.3 (a), existing methods (Tan & Bansal 2019; Lu et al. 2019; Chen et al. 2020) typically introduce a multi-layer perceptron (MLP) multi-label classifier head on top of h[CLS]xh^{x}_{\texttt{[CLS]}}, which…”
    From this paper · §Model
  • LXMERT2019 · cited 9×, 4 in Method
    “As shown in Fig.3 (a), existing methods (Tan & Bansal 2019; Lu et al. 2019; Chen et al. 2020) typically introduce a multi-layer perceptron (MLP) multi-label classifier head on top of h[CLS]xh^{x}_{\texttt{[CLS]}}, which…”
    From this paper · §Model
  • VQA v22016 · cited 7×, 3 in Method
    “Following this success, image+text pretraining models (Lu et al. 2019; Tan & Bansal 2019; Chen et al. 2020; Huang et al. 2020; Li et al. 2020b; Cho et al. 2020; Radford et al. 2021; Zhang et al. 2021) and video+text pret…”
    From this paper · §Related Works
  • ViLBERT2019 · cited 7×, 3 in Method
    “As shown in Fig.3 (a), existing methods (Tan & Bansal 2019; Lu et al. 2019; Chen et al. 2020) typically introduce a multi-layer perceptron (MLP) multi-label classifier head on top of h[CLS]xh^{x}_{\texttt{[CLS]}}, which…”
    From this paper · §Model
Show 14 more
  • T52019 · cited 6×, 3 in Method
    “We introduce VL-T5 and VL-BART based on two pretrained transformer language models: T5Base (Raffel et al. 2020) and BARTBase (Lewis et al. 2020).”
    From this paper · §Model
  • BART2019 · cited 4×, 3 in Method
    “We introduce VL-T5 and VL-BART based on two pretrained transformer language models: T5Base (Raffel et al. 2020) and BARTBase (Lewis et al. 2020).”
    From this paper · §Model
  • BERT2018 · cited 4×, 2 in Method
    “RoI features and bounding box coordinates are encoded with a linear layer, while image ids and region ids are encoded with learned embeddings (Devlin et al. 2019).”
    From this paper · §Model
  • Transformer2017 · cited 3×, 2 in Method
    “We use transformer encoder-decoder architecture (Vaswani et al. 2017) to encode visual and text inputs and generate label text.”
    From this paper · §Model
  • Visual Genome2016 · cited 2×, 2 in Method
    “We represent an input image vv with n=36n{=}36 object regions from a Faster R-CNN (Ren et al. 2015) trained on Visual Genome (Krishna et al. 2016) for object and attribute classification (Anderson et al. 2018).”
    From this paper · §Model
  • COCO Captions2015 · cited 5×, 1 in Method
    “We aggregate pretraining data from MS COCO (Lin et al. 2014; Chen et al. 2015) and Visual Genome (VG; Krishna et al. 2016) images33 3 Existing vision-and-language transformers are trained with different datasets and comp…”
    From this paper · §Pretraining
  • Bottom-Up Top-Down attention2017 · cited 3×, 1 in Method
    “Following this success, image+text pretraining models (Lu et al. 2019; Tan & Bansal 2019; Chen et al. 2020; Huang et al. 2020; Li et al. 2020b; Cho et al. 2020; Radford et al. 2021; Zhang et al. 2021) and video+text pret…”
    From this paper · §Related Works
  • NLVR22018 · cited 3×, 1 in Method
    “Image ids are used to discriminate regions from different images, and is used when multiple images are given to the model (i.e., in NLVR2NLVR^{2} (Suhr et al. 2019), models take two input images).”
    From this paper · §Model
  • MS COCO2014 · cited 2×, 1 in Method
    “We aggregate pretraining data from MS COCO (Lin et al. 2014; Chen et al. 2015) and Visual Genome (VG; Krishna et al. 2016) images33 3 Existing vision-and-language transformers are trained with different datasets and comp…”
    From this paper · §Pretraining
  • Faster R-CNN2015 · cited 1×, 1 in Method
    “We represent an input image vv with n=36n{=}36 object regions from a Faster R-CNN (Ren et al. 2015) trained on Visual Genome (Krishna et al. 2016) for object and attribute classification (Anderson et al. 2018).”
    From this paper · §Model
  • AdamW2017 · cited 1×, 1 in Method
    “We use AdamW (Loshchilov & Hutter 2019) with (β1,β2)=(0.9,0.999)(\beta^{1},\beta^{2})=(0.9,0.999) and learning rate 1e-4 with 5% linear warmup schedule.”
    From this paper · §Pretraining
  • Oscar2020 · cited 4×
    “Following this success, image+text pretraining models (Lu et al. 2019; Tan & Bansal 2019; Chen et al. 2020; Huang et al. 2020; Li et al. 2020b; Cho et al. 2020; Radford et al. 2021; Zhang et al. 2021) and video+text pret…”
    From this paper · §Related Works
  • Unified VLP2019 · cited 2×
    “In this section, we compare our generative architectures VL-T5 and VL-BART on a diverse set of 7 downstream tasks (details in Appendix) with existing vision-and-language pretrained transformers (Tan & Bansal 2019; Lu et…”
    From this paper · §Downstream Tasks and Results
  • CLIP2021 · cited 2×
    “Following this success, image+text pretraining models (Lu et al. 2019; Tan & Bansal 2019; Chen et al. 2020; Huang et al. 2020; Li et al. 2020b; Cho et al. 2020; Radford et al. 2021; Zhang et al. 2021) and video+text pret…”
    From this paper · §Related Works

Led to

  • SimVLM2021 · cited 3×
    “Firstly, we follow Cho et al. 2021 and evaluate model performance on questions with rare answers in the Karpathy-test split.”
    From SimVLM · §Experiments
  • METER2021 · cited 3×, 1 in Method
    “Recently, VL-T5 cho2021unifying and SimVLM wang2021simvlm, on the other hand, advocate the use of a transformer encoder-decoder architecture, where the cross-modal representations are first fed into a decoder and then to…”
    From METER · §The Meter Framework
  • BLIP2022 · cited 3×
    “There have been many attempts to unify various vision and language tasks into a single framework (Zhou et al. 2020; Cho et al. 2021; Wang et al. 2021).”
    From BLIP · §Related Work
  • Flamingo2022 · cited 4×, 1 in Method
    “We accumulate gradients over all datasets, which we found outperforms a “round-robin” approach [17].”
    From Flamingo · §Approach
Abstract

Existing methods for vision-and-language learning typically require designing task-specific architectures and objectives for each task. For example, a multi-label answer classifier for visual question answering, a region scorer for referring expression comprehension, and a language decoder for image captioning, etc. To alleviate these hassles, in this work, we propose a unified framework that learns different tasks in a single architecture with the same language modeling objective, i.e., multimodal conditional text generation, where our models learn to generate labels in text based on the visual and textual inputs. On 7 popular vision-and-language benchmarks, including visual question answering, referring expression comprehension, visual commonsense reasoning, most of which have been previously modeled as discriminative tasks, our generative approach (with a single unified architecture) reaches comparable performance to recent task-specific state-of-the-art vision-and-language models. Moreover, our generative approach shows better generalization ability on questions that have rare answers. Also, we show that our framework allows multi-task learning in a single architecture with a single set of parameters, achieving similar performance to separately optimized single-task models. Our code is publicly available at: https://github.com/j-min/VL-T5