Paper Lineage
Esc
MethodMay 2015arXiv 1505.01121cs.CV

Ask Your Neurons: A Neural-based Approach to Answering Questions about Images

Mateusz Malinowski, Marcus Rohrbach, Mario Fritz

We address a question answering task on real-world images that is set up as a Visual Turing Test. By combining latest advances in image representation and natural language processing, we propose Neural-Image-QA, an end-to-end formulation to this problem for which all parts are trained jointly.

From the abstract

Built on

4 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (4)

  • Seq2Seq2014 · cited 3×, 1 in Method
    “LSTM has been recently shown to be effective in learning a variable-length sequence-to-sequence mapping [5, 28].”
    From this paper · §Approach
  • GoogLeNet (Inception)2014 · cited 2×, 1 in Method
    “In a pilot study, we have found that GoogleNet architecture [11, 29] consistently outperforms the AlexNet architecture [11, 16] as a CNN model for our task and model.”
    From this paper · §Approach
  • Multi-World QA2014 · cited 13×
    “In [20], we present a question answering system based on a semantic parser on a more varied set of human question-answer pairs.”
    From this paper · §Related Work
  • “The task of describing visual content like still images as well as videos has been successfully addressed with a combination of the previous two ideas [5, 12, 31, 32, 37].”
    From this paper · §Related Work

Led to

  • FM-IQA2015 · cited 7×, 2 in Method
    “There are some concurrent and independent works on this topic: [1, 23, 32]. [1] propose a large-scale dataset also based on MS COCO.”
    From FM-IQA · §Related Work
  • Visual Genome2016 · cited 2×
    “The proposed models range from SVM classifiers Antol et al., 2015 and probabilistic inference Malinowski and Fritz, 2014 to recurrent neural networks Gao et al., 2015; Malinowski et al., 2015; Ren et al., 2015a and convo…”
    From Visual Genome · §Related Work
  • VQA v22016 · cited 2×
    “A number of recent works have proposed visual question answering datasets VQA; VisualGenome; fritz; Ren_2015_NIPS; baiduVQA; Madlibs; MovieQA; fsvqa and models MCB; HieCoAtt; NeuralModuleNetworks; dynamic_memory_net_visi…”
    From VQA v2 · §Related Work
Abstract

We address a question answering task on real-world images that is set up as a Visual Turing Test. By combining latest advances in image representation and natural language processing, we propose Neural-Image-QA, an end-to-end formulation to this problem for which all parts are trained jointly. In contrast to previous efforts, we are facing a multi-modal problem where the language output (answer) is conditioned on visual and natural language input (image and question). Our approach Neural-Image-QA doubles the performance of the previous best approach on this problem. We provide additional insights into the problem by analyzing how much information is contained only in the language part for which we provide a new human baseline. To study human consensus, which is related to the ambiguities inherent in this challenging task, we propose two novel metrics and collect additional answers which extends the original DAQUAR dataset to DAQUAR-Consensus.