evaluationIntroduced by VQA · 2015
Open-ended visual question answering
Answer free-form natural-language questions about an image.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that VQA cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Answer questions by reasoning over uncertain scene parses in a probabilistic framework.
Papers using this
- 2016Visual Genome
- 2016VQA v2
- 2017Bottom-Up Top-Down attention
- 2018Bilinear Attention Networks
- 2018NLVR2
- 2018Multi-task hierarchical VL
- 2019ViLBERT
- 2019VisualBERT
- 2019LXMERT
- 2019VL-BERT
- 2019Decoupled box proposals captioning
- 2019Unified VLP
- 2019UNITER
- 201912-in-1
- 2020Grid features for VQA
- 2020X-Linear attention
- 2020VILLA
- 2020VL-BERT meta-analysis
- 2021VinVL
- 2021VL-T5
- 2021ViLT
- 2021Conceptual 12M
- 2021ALBEF
- 2021SimVLM
- 2021VLMo
- 2021METER
- 2021Florence
- 2022BLIP
- 2022CoCa
- 2022VL-BEiT
- 2022BEiT-3