Unifying Vision-and-Language Tasks via Text Generation
Existing methods for vision-and-language learning typically require designing task-specific architectures and objectives for each task. For example, a multi-label answer classifier for visual question answering, a region scorer for referring expression comprehension, and a language decoder for image captioning, etc.
Also cited · not yet reviewed (18)
- UNITER2019 · cited 12×, 4 in Method“As shown in Fig.3 (a), existing methods (Tan & Bansal 2019; Lu et al. 2019; Chen et al. 2020) typically introduce a multi-layer perceptron (MLP) multi-label classifier head on top of h[CLS]xh^{x}_{\texttt{[CLS]}}, which…”From this paper · §Model
- LXMERT2019 · cited 9×, 4 in Method“As shown in Fig.3 (a), existing methods (Tan & Bansal 2019; Lu et al. 2019; Chen et al. 2020) typically introduce a multi-layer perceptron (MLP) multi-label classifier head on top of h[CLS]xh^{x}_{\texttt{[CLS]}}, which…”From this paper · §Model
- VQA v22016 · cited 7×, 3 in Method“Following this success, image+text pretraining models (Lu et al. 2019; Tan & Bansal 2019; Chen et al. 2020; Huang et al. 2020; Li et al. 2020b; Cho et al. 2020; Radford et al. 2021; Zhang et al. 2021) and video+text pret…”From this paper · §Related Works
- ViLBERT2019 · cited 7×, 3 in Method“As shown in Fig.3 (a), existing methods (Tan & Bansal 2019; Lu et al. 2019; Chen et al. 2020) typically introduce a multi-layer perceptron (MLP) multi-label classifier head on top of h[CLS]xh^{x}_{\texttt{[CLS]}}, which…”From this paper · §Model
Show 14 more
- T52019 · cited 6×, 3 in Method“We introduce VL-T5 and VL-BART based on two pretrained transformer language models: T5Base (Raffel et al. 2020) and BARTBase (Lewis et al. 2020).”From this paper · §Model
- BART2019 · cited 4×, 3 in Method“We introduce VL-T5 and VL-BART based on two pretrained transformer language models: T5Base (Raffel et al. 2020) and BARTBase (Lewis et al. 2020).”From this paper · §Model
- BERT2018 · cited 4×, 2 in Method“RoI features and bounding box coordinates are encoded with a linear layer, while image ids and region ids are encoded with learned embeddings (Devlin et al. 2019).”From this paper · §Model
- Transformer2017 · cited 3×, 2 in Method“We use transformer encoder-decoder architecture (Vaswani et al. 2017) to encode visual and text inputs and generate label text.”From this paper · §Model
- Visual Genome2016 · cited 2×, 2 in Method“We represent an input image vv with n=36n{=}36 object regions from a Faster R-CNN (Ren et al. 2015) trained on Visual Genome (Krishna et al. 2016) for object and attribute classification (Anderson et al. 2018).”From this paper · §Model
- COCO Captions2015 · cited 5×, 1 in Method“We aggregate pretraining data from MS COCO (Lin et al. 2014; Chen et al. 2015) and Visual Genome (VG; Krishna et al. 2016) images33 3 Existing vision-and-language transformers are trained with different datasets and comp…”From this paper · §Pretraining
- Bottom-Up Top-Down attention2017 · cited 3×, 1 in Method“Following this success, image+text pretraining models (Lu et al. 2019; Tan & Bansal 2019; Chen et al. 2020; Huang et al. 2020; Li et al. 2020b; Cho et al. 2020; Radford et al. 2021; Zhang et al. 2021) and video+text pret…”From this paper · §Related Works
- NLVR22018 · cited 3×, 1 in Method“Image ids are used to discriminate regions from different images, and is used when multiple images are given to the model (i.e., in NLVR2NLVR^{2} (Suhr et al. 2019), models take two input images).”From this paper · §Model
- MS COCO2014 · cited 2×, 1 in Method“We aggregate pretraining data from MS COCO (Lin et al. 2014; Chen et al. 2015) and Visual Genome (VG; Krishna et al. 2016) images33 3 Existing vision-and-language transformers are trained with different datasets and comp…”From this paper · §Pretraining
- Faster R-CNN2015 · cited 1×, 1 in Method“We represent an input image vv with n=36n{=}36 object regions from a Faster R-CNN (Ren et al. 2015) trained on Visual Genome (Krishna et al. 2016) for object and attribute classification (Anderson et al. 2018).”From this paper · §Model
- AdamW2017 · cited 1×, 1 in Method“We use AdamW (Loshchilov & Hutter 2019) with (β1,β2)=(0.9,0.999)(\beta^{1},\beta^{2})=(0.9,0.999) and learning rate 1e-4 with 5% linear warmup schedule.”From this paper · §Pretraining
- Oscar2020 · cited 4דFollowing this success, image+text pretraining models (Lu et al. 2019; Tan & Bansal 2019; Chen et al. 2020; Huang et al. 2020; Li et al. 2020b; Cho et al. 2020; Radford et al. 2021; Zhang et al. 2021) and video+text pret…”From this paper · §Related Works
- Unified VLP2019 · cited 2דIn this section, we compare our generative architectures VL-T5 and VL-BART on a diverse set of 7 downstream tasks (details in Appendix) with existing vision-and-language pretrained transformers (Tan & Bansal 2019; Lu et…”From this paper · §Downstream Tasks and Results
- CLIP2021 · cited 2דFollowing this success, image+text pretraining models (Lu et al. 2019; Tan & Bansal 2019; Chen et al. 2020; Huang et al. 2020; Li et al. 2020b; Cho et al. 2020; Radford et al. 2021; Zhang et al. 2021) and video+text pret…”From this paper · §Related Works
Led to
- SimVLM2021 · cited 3דFirstly, we follow Cho et al. 2021 and evaluate model performance on questions with rare answers in the Karpathy-test split.”From SimVLM · §Experiments
- METER2021 · cited 3×, 1 in Method“Recently, VL-T5 cho2021unifying and SimVLM wang2021simvlm, on the other hand, advocate the use of a transformer encoder-decoder architecture, where the cross-modal representations are first fed into a decoder and then to…”From METER · §The Meter Framework
- BLIP2022 · cited 3דThere have been many attempts to unify various vision and language tasks into a single framework (Zhou et al. 2020; Cho et al. 2021; Wang et al. 2021).”From BLIP · §Related Work
- Flamingo2022 · cited 4×, 1 in Method“We accumulate gradients over all datasets, which we found outperforms a “round-robin” approach [17].”From Flamingo · §Approach
Abstract
Existing methods for vision-and-language learning typically require designing task-specific architectures and objectives for each task. For example, a multi-label answer classifier for visual question answering, a region scorer for referring expression comprehension, and a language decoder for image captioning, etc. To alleviate these hassles, in this work, we propose a unified framework that learns different tasks in a single architecture with the same language modeling objective, i.e., multimodal conditional text generation, where our models learn to generate labels in text based on the visual and textual inputs. On 7 popular vision-and-language benchmarks, including visual question answering, referring expression comprehension, visual commonsense reasoning, most of which have been previously modeled as discriminative tasks, our generative approach (with a single unified architecture) reaches comparable performance to recent task-specific state-of-the-art vision-and-language models. Moreover, our generative approach shows better generalization ability on questions that have rare answers. Also, we show that our framework allows multi-task learning in a single architecture with a single set of parameters, achieving similar performance to separately optimized single-task models. Our code is publicly available at: https://github.com/j-min/VL-T5