Paper Lineage
Esc
DatasetApr 2015arXiv 1504.00325cs.CV

Microsoft COCO Captions: Data Collection and Evaluation Server

Xinlei Chen, Hao Fang, Tsung-Yi Lin and 4 others

In this paper we describe the Microsoft COCO Caption dataset and evaluation server. When completed, the dataset will contain over one and a half million captions describing over 330,000 images.

From the abstract

Built on

1 paper · 0 verifiedSee as graph

Also cited · not yet reviewed (1)

  • CIDEr2014 · cited 8×
    “Comparing results from different approaches can be difficult since numerous evaluation metrics exist [39, 40, 41, 42].”
    From this paper · §I Introduction

Led to

  • VQA2015 · cited 2×
    “Moreover, these automatic metrics such as BLEU and ROUGE have been found to poorly correlate with human judgement for tasks such as image caption evaluation [6].”
    From VQA · §III VQA Dataset Collection
  • Visual Genome2016 · cited 4×
    “This task is closely related to image captioning Chen et al., 2015; however, results from the two are not directly comparable, as region descriptions are short, incomplete sentences.”
    From Visual Genome · §Experiments
  • “For this task we use the ResNet-50 He et al. 2016 CNN, and train the base model on a combined training set containing 155k images comprised of the MSCOCO Chen et al. 2015 training and validation datasets, and the full Fl…”
    From Constrained beam search captioning · §Experiments
  • nocaps2018 · cited 8×
    “Test images contain previously unseen, or ‘novel’ objects that are drawn from a target distribution (in this case, Open Images [20]) that differs from the source/training distribution (COCO [6]).”
    From nocaps · §Related Work
  • ViLBERT2019 · cited 3×
    “While we address many vision-and-language tasks in Sec. 3.2, we do miss some families of tasks including visually grounded dialog [4, 45], embodied tasks like question answering [7] and instruction following [8], and tex…”
    From ViLBERT · §Related Work
  • VisualBERT2019 · cited 3×, 1 in Method
    “Therefore we reach to a resource of paired data: COCO (Chen et al. 2015) that contains images each paired with 5 independent captions.”
    From VisualBERT · §A Joint Representation Model for Vision and Language
  • Unified VLP2019 · cited 2×
    “We validate VLP in our experiments on both the image captioning and VQA tasks using three challenging benchmarks: COCO Captions [\citeauthoryearChen et al.2015], Flickr30k Captions [\citeauthoryearYoung et al.2014], and…”
    From Unified VLP · §Introduction
  • 12-in-12019 · cited 3×
    “We consider COCOcococaption and Flickr30Kplummer2015flickr30k captioning datasets for this task-group.”
    From 12-in-1 · §Vision-and-Language Tasks
  • Grid features for VQA2020 · cited 3×
    “Through a comprehensive set of experiments, we verified that our observations generalize across different network backbones, different VQA models jiang2018pythia; yu2019deep, different VQA benchmarks antol2015vqa; gurari…”
    From Grid features for VQA · §Introduction
  • VirTex2020 · cited 4×, 1 in Method
    “Training Details: We train on the train2017 split of the COCO Captions dataset [36], which provides 118​K118K images with five captions each.”
    From VirTex · §Method
  • VL-T52021 · cited 5×, 1 in Method
    “We aggregate pretraining data from MS COCO (Lin et al. 2014; Chen et al. 2015) and Visual Genome (VG; Krishna et al. 2016) images33 3 Existing vision-and-language transformers are trained with different datasets and comp…”
    From VL-T5 · §Pretraining
  • ALIGN2021 · cited 1×, 1 in Method
    “Two benchmark datasets are considered: Flickr30K (Plummer et al. 2015) and MSCOCO (Chen et al. 2015).”
    From ALIGN · §Pre-training and Task Transfer
  • Conceptual 12M2021 · cited 4×
    “Smaller but less noisy SBU Captions [61] (1̃M) and COCO Captions [20] (106K) datasets are also of high interest.”
    From Conceptual 12M · §Related Work
  • CoCa2022 · cited 2×
    “We evaluate CoCa on the two standard image-text retrieval benchmarks: MSCOCO [63] and Flickr30K [62].”
    From CoCa · §Experiments
Abstract

In this paper we describe the Microsoft COCO Caption dataset and evaluation server. When completed, the dataset will contain over one and a half million captions describing over 330,000 images. For the training and validation images, five independent human generated captions will be provided. To ensure consistency in evaluation of automatic caption generation algorithms, an evaluation server is used. The evaluation server receives candidate captions and scores them using several popular metrics, including BLEU, METEOR, ROUGE and CIDEr. Instructions for using the evaluation server are provided.