Image captioning
Generating natural-language descriptions of images, and how to evaluate them.
| Concept | Introduced by | Years | Papers |
|---|---|---|---|
| Consensus caption metric (CIDEr) Score a caption by its TF-IDF n-gram agreement with many human references. | CIDEr | 2014–2022 | 14 |
| Region–word visual-semantic alignment Align image regions with sentence fragments, then generate descriptions with a multimodal RNN. | Karpathy visual-semantic alignment | 2014–2015 | 1 |
| Visual word detectors for captioning Detect caption words with multiple-instance learning, then compose sentences with an LM. | Captions to Visual Concepts | 2014–2018 | 2 |
| Image-grounded text generation Train the model to generate the caption conditioned on the image. | — | 2015–2022 | 21 |
| Novel object captioning Describe objects never seen in the caption training data. | Learning like a Child | 2015–2022 | 6 |
| Constrained beam search Force chosen tag words into generated captions at test time, without retraining. | Constrained beam search captioning | 2016–2018 | 1 |
| Object-relation graphs for captioning Encode semantic and spatial relations between objects with graph convolutions. | GCN-LSTM captioning | 2018 | 1 |
| Template-and-slot grounded captioning Generate a sentence template whose slots are filled by detected objects. | Neural Baby Talk | 2018 | 1 |
| Meshed-memory captioning Transformer Memory-augmented region encoding plus mesh connectivity across encoder layers. | Meshed-Memory Transformer | 2019–2022 | 1 |
| Bilinear (X-Linear) attention Second-order bilinear interactions inside attention for captioning. | X-Linear attention | 2020 | 1 |
| Semantic concept tokens for captioning Detector-free captioning that predicts semantic concepts from ViT grid features. | ViTCAP | 2021–2022 | 1 |