Microsoft COCO Captions: Data Collection and Evaluation Server
In this paper we describe the Microsoft COCO Caption dataset and evaluation server. When completed, the dataset will contain over one and a half million captions describing over 330,000 images.
Also cited · not yet reviewed (1)
- CIDEr2014 · cited 8דComparing results from different approaches can be difficult since numerous evaluation metrics exist [39, 40, 41, 42].”From this paper · §I Introduction
Led to
- VQA2015 · cited 2דMoreover, these automatic metrics such as BLEU and ROUGE have been found to poorly correlate with human judgement for tasks such as image caption evaluation [6].”From VQA · §III VQA Dataset Collection
- Visual Genome2016 · cited 4דThis task is closely related to image captioning Chen et al., 2015; however, results from the two are not directly comparable, as region descriptions are short, incomplete sentences.”From Visual Genome · §Experiments
- Constrained beam search captioning2016 · cited 2דFor this task we use the ResNet-50 He et al. 2016 CNN, and train the base model on a combined training set containing 155k images comprised of the MSCOCO Chen et al. 2015 training and validation datasets, and the full Fl…”From Constrained beam search captioning · §Experiments
- nocaps2018 · cited 8דTest images contain previously unseen, or ‘novel’ objects that are drawn from a target distribution (in this case, Open Images [20]) that differs from the source/training distribution (COCO [6]).”From nocaps · §Related Work
- ViLBERT2019 · cited 3דWhile we address many vision-and-language tasks in Sec. 3.2, we do miss some families of tasks including visually grounded dialog [4, 45], embodied tasks like question answering [7] and instruction following [8], and tex…”From ViLBERT · §Related Work
- VisualBERT2019 · cited 3×, 1 in Method“Therefore we reach to a resource of paired data: COCO (Chen et al. 2015) that contains images each paired with 5 independent captions.”From VisualBERT · §A Joint Representation Model for Vision and Language
- Unified VLP2019 · cited 2דWe validate VLP in our experiments on both the image captioning and VQA tasks using three challenging benchmarks: COCO Captions [\citeauthoryearChen et al.2015], Flickr30k Captions [\citeauthoryearYoung et al.2014], and…”From Unified VLP · §Introduction
- 12-in-12019 · cited 3דWe consider COCOcococaption and Flickr30Kplummer2015flickr30k captioning datasets for this task-group.”From 12-in-1 · §Vision-and-Language Tasks
- Grid features for VQA2020 · cited 3דThrough a comprehensive set of experiments, we verified that our observations generalize across different network backbones, different VQA models jiang2018pythia; yu2019deep, different VQA benchmarks antol2015vqa; gurari…”From Grid features for VQA · §Introduction
- VirTex2020 · cited 4×, 1 in Method“Training Details: We train on the train2017 split of the COCO Captions dataset [36], which provides 118K118K images with five captions each.”From VirTex · §Method
- VL-T52021 · cited 5×, 1 in Method“We aggregate pretraining data from MS COCO (Lin et al. 2014; Chen et al. 2015) and Visual Genome (VG; Krishna et al. 2016) images33 3 Existing vision-and-language transformers are trained with different datasets and comp…”From VL-T5 · §Pretraining
- ALIGN2021 · cited 1×, 1 in Method“Two benchmark datasets are considered: Flickr30K (Plummer et al. 2015) and MSCOCO (Chen et al. 2015).”From ALIGN · §Pre-training and Task Transfer
- Conceptual 12M2021 · cited 4דSmaller but less noisy SBU Captions [61] (1̃M) and COCO Captions [20] (106K) datasets are also of high interest.”From Conceptual 12M · §Related Work
- CoCa2022 · cited 2דWe evaluate CoCa on the two standard image-text retrieval benchmarks: MSCOCO [63] and Flickr30K [62].”From CoCa · §Experiments
Abstract
In this paper we describe the Microsoft COCO Caption dataset and evaluation server. When completed, the dataset will contain over one and a half million captions describing over 330,000 images. For the training and validation images, five independent human generated captions will be provided. To ensure consistency in evaluation of automatic caption generation algorithms, an evaluation server is used. The evaluation server receives candidate captions and scores them using several popular metrics, including BLEU, METEOR, ROUGE and CIDEr. Instructions for using the evaluation server are provided.