Paper Lineage
Esc
BenchmarkNov 2014arXiv 1411.5726cs.CV

CIDEr: Consensus-based Image Description Evaluation

Ramakrishna Vedantam, C. Lawrence Zitnick, Devi Parikh

Automatically describing an image with a sentence is a long-standing challenge in computer vision and natural language processing. , there is renewed interest in this area.

From the abstract

Built on

2 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (2)

  • “Recently, generative approaches based on combination of Convolutional and Recurrent Neural Networks DBLP:journals/corr/KarpathyF14; DBLP:journals/corr/ChenZ14a; DBLP:journals/corr/DonahueHGRVSD14; DBLP:journals/corr/Viny…”
    From this paper · §Related Work
  • MS COCO2014 · cited 2×
    “The recently released MS COCO dataset LinECCV14coco contains five sentences for a collection of over 100K images.”
    From this paper · §Related Work

Led to

  • “We additionally report performance on the metrics made available from the MSCOCO captioning challenge,55 5 http://mscoco.org/dataset/#cap2015 which includes scores for BLEU-1 through BLEU-4, METEOR, CIDEr VedantamCORR14,…”
    From Captions to Visual Concepts · §Experimental Results
  • COCO Captions2015 · cited 8×
    “Comparing results from different approaches can be difficult since numerous evaluation metrics exist [39, 40, 41, 42].”
    From COCO Captions · §I Introduction
  • Bottom-Up Top-Down attention2017 · cited 2×, 1 in Method
    “To evaluate caption quality, we use the standard automatic evaluation metrics, namely SPICE spice2016, CIDEr Vedantam2015, METEOR meteor-wmt:2014, ROUGE-L Lin2004 and BLEU Papineni2002.”
    From Bottom-Up Top-Down attention · §Evaluation
  • Neural Baby Talk2018 · cited 2×
    “We report results using the COCO captioning evaluation toolkit lin2014microsoft, which reports the widely used automatic evaluation metrics, BLEU papineni2002bleu, METEOR denkowski2014meteor, CIDEr vedantam2015cider and…”
    From Neural Baby Talk · §Experimental Results
  • nocaps2018 · cited 4×
    “To establish the state-of-the-art on our challenging benchmark, we evaluate two of the best performing existing approaches [27, 2] and report their performance based on well-established evaluation metrics – CIDEr [41] an…”
    From nocaps · §Introduction
  • Meshed-Memory Transformer2019 · cited 2×
    “Following previous works anderson2018bottom, we use the CIDEr-D score as reward, as it well correlates with human judgment vedantam2015cider.”
    From Meshed-Memory Transformer · §Meshed-Memory Transformer
  • Prefix-Tuning2021 · cited 1×, 1 in Method
    “We use the official evaluation script, which reports BLEU Papineni et al. 2002, NIST Belz and Reiter 2006, METEOR Lavie and Agarwal 2007, ROUGE-L Lin 2004, and CIDEr Vedantam et al. 2015.”
    From Prefix-Tuning · §Experimental Setup
  • Conceptual 12M2021 · cited 1×, 1 in Method
    “To measure the performance on image caption generation, we consider the standard metrics BLEU-1,4 [62], ROUGE-L [51], METEOR [10], CIDEr-D [79], and SPICE [4].”
    From Conceptual 12M · §Evaluating Vision-and-Language Pre-Training Data
  • Simple end-to-end captioning2022 · cited 2×, 2 in Method
    “We follow this paradigm to directly optimize the image captioning evaluation metric (CIDEr (Vedantam et al. 2015)) with the widely used self-critical sequence training (SCST) REINFORCE algorithm (Rennie et al. 2017b).”
    From Simple end-to-end captioning · §. Methodology
Abstract

Automatically describing an image with a sentence is a long-standing challenge in computer vision and natural language processing. Due to recent progress in object detection, attribute classification, action recognition, etc., there is renewed interest in this area. However, evaluating the quality of descriptions has proven to be challenging. We propose a novel paradigm for evaluating image descriptions that uses human consensus. This paradigm consists of three main parts: a new triplet-based method of collecting human annotations to measure consensus, a new automated metric (CIDEr) that captures consensus, and two new datasets: PASCAL-50S and ABSTRACT-50S that contain 50 sentences describing each image. Our simple metric captures human judgment of consensus better than existing metrics across sentences generated by various sources. We also evaluate five state-of-the-art image description approaches using this new protocol and provide a benchmark for future comparisons. A version of CIDEr named CIDEr-D is available as a part of MS COCO evaluation server to enable systematic evaluation and benchmarking.