objectiveIntroduced by LXMERT · 2019
QA as a pre-training task
Add image question answering to the vision-language pre-training mix.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that LXMERT cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Represent an image by features of detected object regions instead of a pixel grid.
Papers using this
- 2019VL-BERT
- 2019Unified VLP
- 2020VILLA
- 2020VL-BERT meta-analysis
- 2021ViLT
- 2021ALBEF
- 2021SimVLM