Paper Lineage
Esc
BenchmarkMay 2015arXiv 1505.00468cs.CL

VQA: Visual Question Answering

Aishwarya Agrawal, Jiasen Lu, Stanislaw Antol and 4 others

We propose the task of free-form and open-ended Visual Question Answering (VQA). Given an image and a natural language question about the image, the task is to provide an accurate natural language answer.

From the abstract

Built on

6 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (6)

  • VGG2014 · cited 3×, 3 in Method
    “I: The activations from the last hidden layer of VGGNet [48] are used as 4096-dim image embedding.”
    From this paper · §V VQA Baselines and Methods
  • MS COCO2014 · cited 6×
    “We present a large dataset that contains 204,721 images from the MS COCO dataset [32] and a newly created abstract scene dataset [57, 2] that contains 50,000 scenes.”
    From this paper · §I Introduction
  • Multi-World QA2014 · cited 4×
    “Several recent papers have begun to study visual question answering [19, 36, 50, 3].”
    From this paper · §II Related Work
  • “Related to VQA are the tasks of image tagging [11, 29], image captioning [30, 17, 40, 9, 16, 53, 12, 24, 38, 26] and video captioning [46, 21], where words or sentences are generated to describe visual content.”
    From this paper · §II Related Work
Show 2 more
  • “Related to VQA are the tasks of image tagging [11, 29], image captioning [30, 17, 40, 9, 16, 53, 12, 24, 38, 26] and video captioning [46, 21], where words or sentences are generated to describe visual content.”
    From this paper · §II Related Work
  • COCO Captions2015 · cited 2×
    “Moreover, these automatic metrics such as BLEU and ROUGE have been found to poorly correlate with human judgement for tasks such as image caption evaluation [6].”
    From this paper · §III VQA Dataset Collection

Led to

  • FM-IQA2015 · cited 3×
    “There are some concurrent and independent works on this topic: [1, 23, 32]. [1] propose a large-scale dataset also based on MS COCO.”
    From FM-IQA · §Related Work
  • Visual Genome2016 · cited 16×
    “In addition, inspired by VQA Antol et al., 2015, we also collect an average of 1717 question-answer pairs based on the descriptions for each image.”
    From Visual Genome · §Introduction
  • VQA v22016 · cited 16×, 1 in Method
    “Language and vision problems such as image captioning captioning_msr; captioning_xinlei; captioning_berkeley; captioning_stanford; captioning_google; captioning_toronto; captioning_baidu_ucla and visual question answerin…”
    From VQA v2 · §Introduction
  • NLVR22018 · cited 3×
    “However, commonly used resources for language and vision (Antol et al. 2015; Chen et al. 2016, e.g.,) focus mostly on identification of object properties and few spatial relations (Ferraro et al. 2015; Alikhani and Stone…”
    From NLVR2 · §Introduction
  • “Since the recent successes of deep learning on single modality tasks, multi-modal tasks lying on the intersection of vision and language, such as image captioning mscoco; Young_2014_TACL, visual question answering (VQA)…”
    From Multi-task hierarchical VL · §Introduction
  • ViLBERT2019 · cited 3×
    “We train and evaluate on the VQA 2.0 dataset [3] consisting of 1.1 million questions about COCO images [5] each with 10 answers.”
    From ViLBERT · §Experimental Settings
  • VisualBERT2019 · cited 2×
    “Various tasks such as visual question answering (Antol et al. 2015; Goyal et al. 2017), textual grounding (Kazemzadeh et al. 2014; Plummer et al. 2015), and visual reasoning (Suhr et al. 2019; Zellers et al. 2019) have b…”
    From VisualBERT · §Related Work
  • LXMERT2019 · cited 2×, 1 in Method
    “Besides the two original captioning datasets, we also aggregate three large image question answering (image QA) datasets: VQA v2.0 Antol et al. 2015, GQA balanced version Hudson and Manning 2019, and VG-QA Zhu et al. 201…”
    From LXMERT · §Pre-Training Strategies
  • VL-BERT2019 · cited 2×
    “Meanwhile, for tasks at the intersection of vision and language, such as image captioning (Young et al. 2014; Chen et al. 2015; Sharma et al. 2018), visual question answering (VQA) (Antol et al. 2015; Johnson et al. 2017…”
    From VL-BERT · §Introduction
  • “Other VQA datasets, including VQA1.0 Antol et al. 2015, VQA2.0 Goyal et al. 2017, Visual7W Zhu et al. 2016, COCOQA Ren et al. 2015a, and GQA Hudson and Manning 2019 are completely or partly based on MSCOCO or Visual Geno…”
    From Decoupled box proposals captioning · §Visual Question Answering
  • Grid features for VQA2020 · cited 2×
    “Through a comprehensive set of experiments, we verified that our observations generalize across different network backbones, different VQA models jiang2018pythia; yu2019deep, different VQA benchmarks antol2015vqa; gurari…”
    From Grid features for VQA · §Introduction
  • VL-BERT meta-analysis2020 · cited 1×, 1 in Method
    “COCO (Lin et al. 2014) or VQA (Antol et al. 2015), where the images are strongly-associated with crowdsourced captions or question--answer pairs.”
    From VL-BERT meta-analysis · §Experimental Setup
  • METER2021 · cited 3×, 1 in Method
    “Vision-and-language (VL) tasks, such as visual question answering (VQA) antol2015vqa and image-text retrieval lin2014microsoft; plummer2015flickr30k, require an AI system to comprehend both the input image and text conte…”
    From METER · §Introduction
Abstract

We propose the task of free-form and open-ended Visual Question Answering (VQA). Given an image and a natural language question about the image, the task is to provide an accurate natural language answer. Mirroring real-world scenarios, such as helping the visually impaired, both the questions and answers are open-ended. Visual questions selectively target different areas of an image, including background details and underlying context. As a result, a system that succeeds at VQA typically needs a more detailed understanding of the image and complex reasoning than a system producing generic image captions. Moreover, VQA is amenable to automatic evaluation, since many open-ended answers contain only a few words or a closed set of answers that can be provided in a multiple-choice format. We provide a dataset containing ~0.25M images, ~0.76M questions, and ~10M answers (www.visualqa.org), and discuss the information it provides. Numerous baselines and methods for VQA are provided and compared with human performance. Our VQA demo is available on CloudCV (http://cloudcv.org/vqa).