From Captions to Visual Concepts and Back
This paper presents a novel approach for automatically generating image descriptions: visual detectors, language models, and multimodal similarity models learnt directly from a dataset of image captions. We use multiple instance learning to train visual detectors for words that commonly occur in captions, including many different parts of speech such as nouns, verbs, and adjectives.
Also cited · not yet reviewed (4)
- VGG2014 · cited 5דNext, we featurize each of these regions using rich convolutional neural network (CNN) features, fine-tuned on our training data krizhevskyNIPS12; simonyan14very.”From this paper · §Introduction
- MS COCO2014 · cited 2דThe evaluation was performed on the challenging Microsoft COCO dataset linECCV14; capeval2015 containing complex images with multiple objects.”From this paper · §Introduction
- CIDEr2014 · cited 2דWe additionally report performance on the metrics made available from the MSCOCO captioning challenge,55 5 http://mscoco.org/dataset/#cap2015 which includes scores for BLEU-1 through BLEU-4, METEOR, CIDEr VedantamCORR14,…”From this paper · §Experimental Results
- Karpathy visual-semantic alignment2014 · cited 2דAlso related are several contemporaneous papers mao2014explain; Vinyals2014b; Chen2014; Karpathy2014b; Donahue2014a; Xu2015; Lebret2015.”From this paper · §Related Work
Led to
- Learning like a Child2015 · cited 3דVery recent works of image captioning includes [39, 25, 24, 55, 10, 15, 7, 32, 37, 26, 57, 35].”From Learning like a Child · §Related Work
- VQA2015 · cited 2דRelated to VQA are the tasks of image tagging [11, 29], image captioning [30, 17, 40, 9, 16, 53, 12, 24, 38, 26] and video captioning [46, 21], where words or sentences are generated to describe visual content.”From VQA · §II Related Work
- Flickr30k Entities2015 · cited 6דWe build on the Flickr30k dataset (Young et al., 2014), a popular benchmark for caption generation and retrieval that has been used, among others, by Chen and Zitnick, 2015; Donahue et al., 2015; Fang et al., 2015; Gong…”From Flickr30k Entities · §Introduction
- Constrained beam search captioning2016 · cited 2דMore generally, the effectiveness of incorporating semantic attributes (i.e., image tags) into caption model training for in-domain data has been established by several works Fang et al. 2015; Wu et al. 2016; Elliot and…”From Constrained beam search captioning · §Related Work
- Neural Baby Talk2018 · cited 2דWhile there are many recent extensions of this basic idea to include attention Xu2015show; fang2015captions; you2016image; yang2016encode; lu2016knowing, it is well-understood that models still lack visual grounding (i.e…”From Neural Baby Talk · §Introduction
- VinVL2021 · cited 2×, 2 in Method“Different from the binary contrastive loss used in Oscar [21], the proposed 3-way Contrastive Loss to effectively optimize the training objectives used for VQA [41] and text-image matching [6]77 7 [6] uses a deep-learnin…”From VinVL · §Oscar+ Pre-training
Abstract
This paper presents a novel approach for automatically generating image descriptions: visual detectors, language models, and multimodal similarity models learnt directly from a dataset of image captions. We use multiple instance learning to train visual detectors for words that commonly occur in captions, including many different parts of speech such as nouns, verbs, and adjectives. The word detector outputs serve as conditional inputs to a maximum-entropy language model. The language model learns from a set of over 400,000 image descriptions to capture the statistics of word usage. We capture global semantics by re-ranking caption candidates using sentence-level features and a deep multimodal similarity model. Our system is state-of-the-art on the official Microsoft COCO benchmark, producing a BLEU-4 score of 29.1%. When human judges compare the system captions to ones written by other people on our held-out test set, the system captions have equal or better quality 34% of the time.