architectureIntroduced by Ask Your Neurons · 2015
End-to-end neural VQA (CNN + LSTM)
Encode the image with a CNN and the question with an LSTM, and train everything jointly.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that Ask Your Neurons cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Answer questions by reasoning over uncertain scene parses in a probabilistic framework.
Papers using this
- 2015Image QA models & data
- 2015FM-IQA