Paper Lineage
Esc

Image captioning

Generating natural-language descriptions of images, and how to evaluate them.

Concepts in Image captioning
ConceptIntroduced byYearsPapers
Consensus caption metric (CIDEr)
Score a caption by its TF-IDF n-gram agreement with many human references.
CIDEr2014–202214
Region–word visual-semantic alignment
Align image regions with sentence fragments, then generate descriptions with a multimodal RNN.
Karpathy visual-semantic alignment2014–20151
Visual word detectors for captioning
Detect caption words with multiple-instance learning, then compose sentences with an LM.
Captions to Visual Concepts2014–20182
Image-grounded text generation
Train the model to generate the caption conditioned on the image.
—2015–202221
Novel object captioning
Describe objects never seen in the caption training data.
Learning like a Child2015–20226
Constrained beam search
Force chosen tag words into generated captions at test time, without retraining.
Constrained beam search captioning2016–20181
Object-relation graphs for captioning
Encode semantic and spatial relations between objects with graph convolutions.
GCN-LSTM captioning20181
Template-and-slot grounded captioning
Generate a sentence template whose slots are filled by detected objects.
Neural Baby Talk20181
Meshed-memory captioning Transformer
Memory-augmented region encoding plus mesh connectivity across encoder layers.
Meshed-Memory Transformer2019–20221
Bilinear (X-Linear) attention
Second-order bilinear interactions inside attention for captioning.
X-Linear attention20201
Semantic concept tokens for captioning
Detector-free captioning that predicts semantic concepts from ViT grid features.
ViTCAP2021–20221