Visual question answering & reasoning
Answering questions and reasoning about images in natural language.
| Concept | Introduced by | Years | Papers |
|---|---|---|---|
| Symbolic reasoning for visual QA Answer questions by reasoning over uncertain scene parses in a probabilistic framework. | Multi-World QA | 2014–2015 | 1 |
| End-to-end neural VQA (CNN + LSTM) Encode the image with a CNN and the question with an LSTM, and train everything jointly. | Ask Your Neurons | 2015 | 2 |
| Generating QA pairs from captions Turn existing image descriptions into question–answer training data automatically. | Image QA models & data | 2015–2019 | 1 |
| Open-ended visual question answering Answer free-form natural-language questions about an image. | VQA | 2015–2022 | 31 |
| Balanced VQA against language priors Pair every question with images that flip the answer, so models must actually look. | VQA v2 | 2016–2022 | 6 |
| Bilinear attention networks Attend over all question-word × image-region pairs with low-rank bilinear pooling. | Bilinear Attention Networks | 2018–2020 | 2 |
| Grounded visual reasoning Decide whether a statement is true of a pair of photos, which requires compositional reasoning. | NLVR2 | 2018–2022 | 12 |