objective
Image-grounded text generation
Train the model to generate the caption conditioned on the image.
Drafted by AI · not yet reviewed
How this idea evolved
No earlier ideas recorded for this concept yet.
Papers using this
- 2015Learning like a Child
- 2015FM-IQA
- 2016Constrained beam search captioning
- 2017Bottom-Up Top-Down attention
- 2018Neural Baby Talk
- 2018GCN-LSTM captioning
- 2018nocaps
- 2019Decoupled box proposals captioning
- 2019Unified VLP
- 2019Localized Narratives
- 2019Meshed-Memory Transformer
- 2020Grid features for VQA
- 2020X-Linear attention
- 2021VL-T5
- 2021Conceptual 12M
- 2021SimVLM
- 2021ViTCAP
- 2022BLIP
- 2022Simple end-to-end captioning
- 2022CoCa
- 2022BEiT-3