evaluationIntroduced by CIDEr · 2014
Consensus caption metric (CIDEr)
Score a caption by its TF-IDF n-gram agreement with many human references.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that CIDEr cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Align image regions with sentence fragments, then generate descriptions with a multimodal RNN.
Papers using this
- 2014Captions to Visual Concepts
- 2015COCO Captions
- 2016Visual Genome
- 2017Bottom-Up Top-Down attention
- 2018Neural Baby Talk
- 2018GCN-LSTM captioning
- 2018nocaps
- 2020X-Linear attention
- 2021Prefix-Tuning
- 2021Conceptual 12M
- 2021SimVLM
- 2021ViTCAP
- 2022BLIP
- 2022Simple end-to-end captioning