dataIntroduced by VQA v2 · 2016
Balanced VQA against language priors
Pair every question with images that flip the answer, so models must actually look.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that VQA v2 cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Answer free-form natural-language questions about an image.
Encode the image with a CNN and the question with an LSTM, and train everything jointly.
Turn existing image descriptions into question–answer training data automatically.