Paper Lineage
Esc
MethodMay 2015arXiv 1505.02074cs.LG

Exploring Models and Data for Image Question Answering

Mengye Ren, Ryan Kiros, Richard Zemel

This work aims to address the problem of image-based question-answering (QA) with new models and datasets. In our work, we propose to use neural networks and visual semantic embeddings, without intermediate stages such as object detection and image segmentation, to predict answers to simple questions about images.

From the abstract

Built on

0 papers · 0 verifiedSee as graph

Image QA models & data has no earlier papers in this dataset.

Led to

  • FM-IQA2015 · cited 7×, 2 in Method
    “There are some concurrent and independent works on this topic: [1, 23, 32]. [1] propose a large-scale dataset also based on MS COCO.”
    From FM-IQA · §Related Work
  • Visual Genome2016 · cited 6×
    “Most new datasets Yu et al., 2015; Ren et al., 2015a; Antol et al., 2015; Gao et al., 2015 have collected QA pairs on MS-COCO images, either generated automatically by NLP tools Ren et al., 2015a or written by human work…”
    From Visual Genome · §Related Work
  • VQA v22016 · cited 2×
    “A number of recent works have proposed visual question answering datasets VQA; VisualGenome; fritz; Ren_2015_NIPS; baiduVQA; Madlibs; MovieQA; fsvqa and models MCB; HieCoAtt; NeuralModuleNetworks; dynamic_memory_net_visi…”
    From VQA v2 · §Related Work
Abstract

This work aims to address the problem of image-based question-answering (QA) with new models and datasets. In our work, we propose to use neural networks and visual semantic embeddings, without intermediate stages such as object detection and image segmentation, to predict answers to simple questions about images. Our model performs 1.8 times better than the only published results on an existing image QA dataset. We also present a question generation algorithm that converts image descriptions, which are widely available, into QA form. We used this algorithm to produce an order-of-magnitude larger dataset, with more evenly distributed answers. A suite of baseline results on this new dataset are also presented.