A Multi-World Approach to Question Answering about Real-World Scenes based on Uncertain Input
We propose a method for automatically answering questions about images by bringing together recent advances from natural language processing and computer vision. We combine discrete reasoning with uncertain predictions by a multi-world approach that represents uncertainty about the perceived world in a bayesian framework.
Multi-World QA has no earlier papers in this dataset.
Led to
- VQA2015 · cited 4דSeveral recent papers have begun to study visual question answering [19, 36, 50, 3].”From VQA · §II Related Work
- Ask Your Neurons2015 · cited 13דIn [20], we present a question answering system based on a semantic parser on a more varied set of human question-answer pairs.”From Ask Your Neurons · §Related Work
- FM-IQA2015 · cited 3דThere has been recent effort on the visual question answering task [9, 2, 22, 37].”From FM-IQA · §Related Work
- Visual Genome2016 · cited 9דVisual question answering (QA) has been recently proposed as a proxy task of evaluating a computer vision system’s ability to understand an image beyond object recognition Geman et al., 2015; Malinowski and Fritz, 2014.”From Visual Genome · §Related Work
- VQA v22016 · cited 2דA number of recent works have proposed visual question answering datasets VQA; VisualGenome; fritz; Ren_2015_NIPS; baiduVQA; Madlibs; MovieQA; fsvqa and models MCB; HieCoAtt; NeuralModuleNetworks; dynamic_memory_net_visi…”From VQA v2 · §Related Work
Abstract
We propose a method for automatically answering questions about images by bringing together recent advances from natural language processing and computer vision. We combine discrete reasoning with uncertain predictions by a multi-world approach that represents uncertainty about the perceived world in a bayesian framework. Our approach can handle human questions of high complexity about realistic scenes and replies with range of answer like counts, object classes, instances and lists of them. The system is directly trained from question-answer pairs. We establish a first benchmark for this task that can be seen as a modern attempt at a visual turing test.