Paper Lineage
Esc
MethodDec 2014arXiv 1412.2306cs.CV

Deep Visual-Semantic Alignments for Generating Image Descriptions

Andrej Karpathy, Li Fei-Fei

We present a model that generates natural language descriptions of images and their regions. Our approach leverages datasets of images and their sentence descriptions to learn about the inter-modal correspondences between language and visual data.

From the abstract

Built on

4 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (4)

  • Deep Fragment Embeddings2014 · cited 10×, 5 in Method
    “We build on the approach of Karpathy et al. defrag, who learn to ground dependency tree relations to image regions with a ranking objective.”
    From this paper · §Our Model
  • VGG2014 · cited 3×
    “We list their performance with a CNN that is equivalent in power (AlexNet krizhevsky2012imagenet) to the one used in this work, though similar to vinyals2014show they outperform our model with a more powerful CNN (VGGNet…”
    From this paper · §Experiments
  • MS COCO2014 · cited 2×
    “The second, practical challenge is that datasets of image captions are available in large quantities on the internet hodosh2013framing; flickr30k; coco, but these descriptions multiplex mentions of several entities whose…”
    From this paper · §Introduction
  • GoogLeNet (Inception)2014 · cited 2×
    “We list their performance with a CNN that is equivalent in power (AlexNet krizhevsky2012imagenet) to the one used in this work, though similar to vinyals2014show they outperform our model with a more powerful CNN (VGGNet…”
    From this paper · §Experiments

Led to

  • “Also related are several contemporaneous papers mao2014explain; Vinyals2014b; Chen2014; Karpathy2014b; Donahue2014a; Xu2015; Lebret2015.”
    From Captions to Visual Concepts · §Related Work
  • CIDEr2014 · cited 3×
    “Recently, generative approaches based on combination of Convolutional and Recurrent Neural Networks DBLP:journals/corr/KarpathyF14; DBLP:journals/corr/ChenZ14a; DBLP:journals/corr/DonahueHGRVSD14; DBLP:journals/corr/Viny…”
    From CIDEr · §Related Work
  • VQA2015 · cited 2×
    “Related to VQA are the tasks of image tagging [11, 29], image captioning [30, 17, 40, 9, 16, 53, 12, 24, 38, 26] and video captioning [46, 21], where words or sentences are generated to describe visual content.”
    From VQA · §II Related Work
  • Ask Your Neurons2015 · cited 3×
    “The task of describing visual content like still images as well as videos has been successfully addressed with a combination of the previous two ideas [5, 12, 31, 32, 37].”
    From Ask Your Neurons · §Related Work
  • Flickr30k Entities2015 · cited 9×
    “We build on the Flickr30k dataset (Young et al., 2014), a popular benchmark for caption generation and retrieval that has been used, among others, by Chen and Zitnick, 2015; Donahue et al., 2015; Fang et al., 2015; Gong…”
    From Flickr30k Entities · §Introduction
  • “Recently, several approaches based on RNNs emerged, generating captions via a learned joint image-text embedding [13, 11, 36, 21].”
    From BookCorpus (books & movies) · §Related Work
  • Visual Genome2016 · cited 9×
    “We train one of the top 16 state-of-the-art image caption generator Karpathy and Fei-Fei, 2014 on (1) our dataset to generate region descriptions and on (2) Flickr30K Young et al., 2014 to generate sentence descriptions.”
    From Visual Genome · §Experiments
  • Constrained beam search captioning2016 · cited 2×, 1 in Method
    “Our approach to out-of-domain image captioning could be applied to any existing CNN-RNN captioning model that can be decoding using beam search, e.g., Donahue et al. 2015; Mao et al. 2015; Karpathy and Fei-Fei 2015; Viny…”
    From Constrained beam search captioning · §Approach
  • Visual N-Grams2016 · cited 5×
    “Relating Images and Captions: Additional Results As an addition to the image and caption retrieval results on COCO-5K and Flickr-30K presented in the paper, we also provide retrieval results on the COCO-1K dataset, a tes…”
    From Visual N-Grams · §Introduction
  • Neural Baby Talk2018 · cited 3×
    “For standard image captioning, we use splits from Karpathy et al. karpathy2015deep on COCO/Flickr30k.”
    From Neural Baby Talk · §Experimental Results
  • “Following the standard procedure Karpathy_2015_CVPR, we use the 1,000 val images and the 1,000 or 5,000 test images, which are selected from the original 40,504 val images.”
    From Multi-task hierarchical VL · §Experiments
  • Meshed-Memory Transformer2019 · cited 2×
    “We follow the splits provided by Karpathy et al. karpathy2015deep, where 5 0005\,000 images are used for validation, 5 0005\,000 for testing and the rest for training.”
    From Meshed-Memory Transformer · §Experiments
  • ViLT2021 · cited 2×
    “We evaluate ViLT on two widely explored types of vision-and-language downstream tasks: for classification, we use VQAv2 (Goyal et al. 2017) and NLVR2 (Suhr et al. 2018), and for retrieval, we use MSCOCO and Flickr30K (F3…”
    From ViLT · §Experiments
  • Simple end-to-end captioning2022 · cited 1×, 1 in Method
    “We follow the standard Karpathy’s split (Karpathy and Fei-Fei 2015) to split 113.2k/5k/5k and 29.8k/1k/1k images for train/val/test, respectively.”
    From Simple end-to-end captioning · §. Experiment Setup
  • BEiT-32022 · cited 2×
    “We use the COCO [31] benchmark, finetune and evaluate the model on Karpathy split [24].”
    From BEiT-3 · §Experiments on Vision and Vision-Language Tasks
Abstract

We present a model that generates natural language descriptions of images and their regions. Our approach leverages datasets of images and their sentence descriptions to learn about the inter-modal correspondences between language and visual data. Our alignment model is based on a novel combination of Convolutional Neural Networks over image regions, bidirectional Recurrent Neural Networks over sentences, and a structured objective that aligns the two modalities through a multimodal embedding. We then describe a Multimodal Recurrent Neural Network architecture that uses the inferred alignments to learn to generate novel descriptions of image regions. We demonstrate that our alignment model produces state of the art results in retrieval experiments on Flickr8K, Flickr30K and MSCOCO datasets. We then show that the generated descriptions significantly outperform retrieval baselines on both full images and on a new dataset of region-level annotations.