Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Problems at the intersection of vision and language are of significant importance both as challenging research questions and for the rich set of applications they enable. However, inherent structure in our world and bias in our language tend to be a simpler signal for learning than visual modalities, resulting in models that ignore visual information, leading to an inflated sense of their capability.
Also cited · not yet reviewed (8)
- VQA2015 · cited 16×, 1 in Method“Language and vision problems such as image captioning captioning_msr; captioning_xinlei; captioning_berkeley; captioning_stanford; captioning_google; captioning_toronto; captioning_baidu_ucla and visual question answerin…”From this paper · §Introduction
- VGG2014 · cited 3×, 1 in Method“It should be noted that MCB uses image features from a more powerful CNN architecture ResNet ResNet while the previous two models use image features from VGGNet Simonyan15.”From this paper · §Benchmarking Existing VQA Models
- ResNet2015 · cited 1×, 1 in Method“It should be noted that MCB uses image features from a more powerful CNN architecture ResNet ResNet while the previous two models use image features from VGGNet Simonyan15.”From this paper · §Benchmarking Existing VQA Models
- MS COCO2014 · cited 2דOur complete balanced dataset contains approximately 1.1 Million (image, question) pairs – almost double the size of the VQA VQA dataset – with approximately 13 Million associated answers on the ∼\sim200k images from COC…”From this paper · §Introduction
Show 4 more
- Multi-World QA2014 · cited 2דA number of recent works have proposed visual question answering datasets VQA; VisualGenome; fritz; Ren_2015_NIPS; baiduVQA; Madlibs; MovieQA; fsvqa and models MCB; HieCoAtt; NeuralModuleNetworks; dynamic_memory_net_visi…”From this paper · §Related Work
- Ask Your Neurons2015 · cited 2דA number of recent works have proposed visual question answering datasets VQA; VisualGenome; fritz; Ren_2015_NIPS; baiduVQA; Madlibs; MovieQA; fsvqa and models MCB; HieCoAtt; NeuralModuleNetworks; dynamic_memory_net_visi…”From this paper · §Related Work
- Image QA models & data2015 · cited 2דA number of recent works have proposed visual question answering datasets VQA; VisualGenome; fritz; Ren_2015_NIPS; baiduVQA; Madlibs; MovieQA; fsvqa and models MCB; HieCoAtt; NeuralModuleNetworks; dynamic_memory_net_visi…”From this paper · §Related Work
- FM-IQA2015 · cited 2דLanguage and vision problems such as image captioning captioning_msr; captioning_xinlei; captioning_berkeley; captioning_stanford; captioning_google; captioning_toronto; captioning_baidu_ucla and visual question answerin…”From this paper · §Introduction
Led to
- Bottom-Up Top-Down attention2017 · cited 2דProblems combining image and language understanding such as image captioning Chen2015 and visual question answering (VQA) balanced_vqa_v2 continue to inspire considerable research at the boundary of computer vision and n…”From Bottom-Up Top-Down attention · §Introduction
- Bilinear Attention Networks2018 · cited 2דFinally, we validate our proposed method on a large and highly-competitive dataset, VQA 2.0 [8].”From Bilinear Attention Networks · §Introduction
- NLVR22018 · cited 2דThis relatively simple reasoning, together with biases in the data, removes much of the need to consider language compositionality Goyal et al. 2017.”From NLVR2 · §Introduction
- Multi-task hierarchical VL2018 · cited 4דSince the recent successes of deep learning on single modality tasks, multi-modal tasks lying on the intersection of vision and language, such as image captioning mscoco; Young_2014_TACL, visual question answering (VQA)…”From Multi-task hierarchical VL · §Introduction
- VisualBERT2019 · cited 4דWe evaluate VisualBERT on four different types of vision-and-language applications: (1) Visual Question Answering (VQA 2.0) (Goyal et al. 2017), (2) Visual Commonsense Reasoning (VCR) (Zellers et al. 2019), (3) Natural L…”From VisualBERT · §Experiment
- LXMERT2019 · cited 1×, 1 in Method“We use three datasets for evaluating our LXMERT framework: VQA v2.0 dataset Goyal et al. 2017, GQA Hudson and Manning 2019, and NLVR2NLVR^{2}.”From LXMERT · §Experimental Setup and Results
- VL-BERT2019 · cited 3דHere we conduct experiments on the widely-used VQA v2.0 dataset (Goyal et al. 2017), which is built based on the COCO (Lin et al. 2014) images.”From VL-BERT · §Experiment
- Decoupled box proposals captioning2019 · cited 2×, 1 in Method“Other VQA datasets, including VQA1.0 Antol et al. 2015, VQA2.0 Goyal et al. 2017, Visual7W Zhu et al. 2016, COCOQA Ren et al. 2015a, and GQA Hudson and Manning 2019 are completely or partly based on MSCOCO or Visual Geno…”From Decoupled box proposals captioning · §Visual Question Answering
- Grid features for VQA2020 · cited 4דIt is also worth noting that while region features are effective on benchmarks like VQA antol2015vqa; goyal2017making and COCO captions chen2015microsoft, for benchmarks that diagnose a model’s reasoning abilities when a…”From Grid features for VQA · §Related Work
- VILLA2020 · cited 2דWhen finetuned on downstream tasks, these pre-trained models have achieved state-of-the-art performance across diverse V+L tasks, such as Visual Question Answering (VQA) [4, 17], Visual Commonsense Reasoning (VCR) [81],…”From VILLA · §Introduction
- VL-BERT meta-analysis2020 · cited 2×, 1 in Method“We consider the most common tasks used to evaluate V&L BERTs, spanning four groups: vocab-based VQA (Goyal et al. 2017; Hudson and Manning 2019), image–text retrieval (Lin et al. 2014; Plummer et al. 2015), referring exp…”From VL-BERT meta-analysis · §Experimental Setup
- VinVL2021 · cited 3×, 1 in Method“We build our pre-training corpus based on three types of existing vision and VL datasets: (1) image captioning datasets with human-annotated captions as 𝒘\boldsymbol{w} and machine-generated 55 5 We use the same model t…”From VinVL · §Oscar+ Pre-training
- VL-T52021 · cited 7×, 3 in Method“Following this success, image+text pretraining models (Lu et al. 2019; Tan & Bansal 2019; Chen et al. 2020; Huang et al. 2020; Li et al. 2020b; Cho et al. 2020; Radford et al. 2021; Zhang et al. 2021) and video+text pret…”From VL-T5 · §Related Works
- Conceptual 12M2021 · cited 2דSecond, many of the small-sized datasets share the same, limited visual domain; COCO-Captions [20], Visual Genome [44], and VQA2 [27] are (mostly) based on several hundreds thousand of COCO images [52].”From Conceptual 12M · §Introduction
- SimVLM2021 · cited 2דA line of work (Tan & Bansal 2019; Lu et al. 2019; Li et al. 2019; Chen et al. 2020b; Li et al. 2020; Su et al. 2020; Zhang et al. 2021) has explored vision-language pretraining (VLP) that learns a joint representation o…”From SimVLM · §Introduction
- VLMo2021 · cited 2דWe train and evaluate the model on VQA 2.0 dataset [15].”From VLMo · §Experiments
- Florence2021 · cited 2×, 1 in Method“To evaluate the performance, we fine-tune the pre-trained model on the challenging VQA (Goyal et al. 2017) task, which is to answer a question based on the image context.”From Florence · §Experiments
- VL-BEiT2022 · cited 2דFollowing previous work [16, 41], we use VQA 2.0 dataset [10], and formulate the task as a classification problem to choose the answer from 3,1293,129 most frequent answers.”From VL-BEiT · §Experiments
- BEiT-32022 · cited 2דFollowing previous work [2, 65, 26], we conduct finetuning experiments on the VQA v2.0 dataset [19] and formulate the task as a classification problem.”From BEiT-3 · §Experiments on Vision and Vision-Language Tasks
Abstract
Problems at the intersection of vision and language are of significant importance both as challenging research questions and for the rich set of applications they enable. However, inherent structure in our world and bias in our language tend to be a simpler signal for learning than visual modalities, resulting in models that ignore visual information, leading to an inflated sense of their capability. We propose to counter these language priors for the task of Visual Question Answering (VQA) and make vision (the V in VQA) matter! Specifically, we balance the popular VQA dataset by collecting complementary images such that every question in our balanced dataset is associated with not just a single image, but rather a pair of similar images that result in two different answers to the question. Our dataset is by construction more balanced than the original VQA dataset and has approximately twice the number of image-question pairs. Our complete balanced dataset is publicly available at www.visualqa.org as part of the 2nd iteration of the Visual Question Answering Dataset and Challenge (VQA v2.0). We further benchmark a number of state-of-art VQA models on our balanced dataset. All models perform significantly worse on our balanced dataset, suggesting that these models have indeed learned to exploit language priors. This finding provides the first concrete empirical evidence for what seems to be a qualitative sense among practitioners. Finally, our data collection protocol for identifying complementary images enables us to develop a novel interpretable model, which in addition to providing an answer to the given (image, question) pair, also provides a counter-example based explanation. Specifically, it identifies an image that is similar to the original image, but it believes has a different answer to the same question. This can help in building trust for machines among their users.