Paper Lineage
Esc
DatasetMay 2014arXiv 1405.0312cs.CV

Microsoft COCO: Common Objects in Context

Tsung-Yi Lin, Michael Maire, Serge Belongie and 7 others

We present a new dataset with the goal of advancing the state-of-the-art in object recognition by placing the question of object recognition in the context of the broader question of scene understanding. This is achieved by gathering images of complex everyday scenes containing common objects in their natural context.

From the abstract

Built on

0 papers · 0 verifiedSee as graph

MS COCO has no earlier papers in this dataset.

Led to

  • “The evaluation was performed on the challenging Microsoft COCO dataset linECCV14; capeval2015 containing complex images with multiple objects.”
    From Captions to Visual Concepts · §Introduction
  • CIDEr2014 · cited 2×
    “The recently released MS COCO dataset LinECCV14coco contains five sentences for a collection of over 100K images.”
    From CIDEr · §Related Work
  • “The second, practical challenge is that datasets of image captions are available in large quantities on the internet hodosh2013framing; flickr30k; coco, but these descriptions multiplex mentions of several entities whose…”
    From Karpathy visual-semantic alignment · §Introduction
  • Learning like a Child2015 · cited 2×
    “The first two datasets are derived from the MS-COCO dataset [34].”
    From Learning like a Child · §Introduction
  • VQA2015 · cited 6×
    “We present a large dataset that contains 204,721 images from the MS COCO dataset [32] and a newly created abstract scene dataset [57, 2] that contains 50,000 scenes.”
    From VQA · §I Introduction
  • FM-IQA2015 · cited 3×
    “The large-scale image datasets with sentence annotations (e.g., [21, 43, 11]) play a crucial role in this progress.”
    From FM-IQA · §Introduction
  • Visual Genome2016 · cited 15×
    “MS-COCO Lin et al., 2014 recently released its dataset, with over 328,000328,000 images with sentence descriptions and segmentations of 9191 object categories.”
    From Visual Genome · §Related Work
  • ResNeXt2016 · cited 3×
    “This paper further evaluates ResNeXt on a larger ImageNet-5K set and the COCO object detection dataset Lin2014, showing consistently better accuracy than its ResNet counterparts.”
    From ResNeXt · §Introduction
  • “Recently, models incorporating recurrent neural networks (RNNs) have demonstrated promising results on this challenging task Vinyals et al. 2015; Fang et al. 2015; Devlin et al. 2015, leveraging new benchmark datasets su…”
    From Constrained beam search captioning · §Introduction
  • VQA v22016 · cited 2×
    “Our complete balanced dataset contains approximately 1.1 Million (image, question) pairs – almost double the size of the VQA VQA dataset – with approximately 13 Million associated answers on the ∼\sim200k images from COC…”
    From VQA v2 · §Introduction
  • “We evaluate on the two most popular datasets: COCO [26] and PASCAL VOC [13].”
    From JFT-300M (unreasonable effectiveness) · §Experiments
  • “As approximately 51K Visual Genome images are also found in the MSCOCO captions dataset Lin2014, we are careful to avoid contamination of our MSCOCO validation and test sets.”
    From Bottom-Up Top-Down attention · §Evaluation
  • Neural Baby Talk2018 · cited 5×, 2 in Method
    “COCO lin2014microsoft) do not contain phrase grounding annotations, while some datasets do (e.g.”
    From Neural Baby Talk · §Method
  • “For example, we observe improvements over the state-of-the-art for image classification and object detection, where we obtain a single-crop, top-1 accuracy of 85.4% on the ImageNet-1k image-classification dataset and 45.…”
    From Instagram hashtag pre-training · §Introduction
  • MnasNet2018 · cited 2×
    “We apply our proposed approach to ImageNet classification imagenet15 and COCO object detection coco14.”
    From MnasNet · §Introduction
  • NLVR22018 · cited 2×
    “We detect the objects in the images using a Mask R-CNN model He et al. 2017; Girshick et al. 2018 pre-trained on the COCO detection task Lin et al. 2014.”
    From NLVR2 · §Evaluation Systems
  • “Since the recent successes of deep learning on single modality tasks, multi-modal tasks lying on the intersection of vision and language, such as image captioning mscoco; Young_2014_TACL, visual question answering (VQA)…”
    From Multi-task hierarchical VL · §Introduction
  • LVIS2019 · cited 2×
    “Quality is important for future research because relatively coarse masks, such as those in the COCO dataset Lin2014, limit the ability to differentiate algorithm-predicted mask quality beyond a certain, coarse point.”
    From LVIS · §Introduction
  • LXMERT2019 · cited 2×, 1 in Method
    “As shown in Table. 1, we aggregate pre-training data from five vision-and-language datasets whose images come from MS COCO Lin et al. 2014 or Visual Genome Krishna et al. 2017.”
    From LXMERT · §Pre-Training Strategies
  • VL-BERT2019 · cited 2×
    “Here we conduct experiments on the widely-used VQA v2.0 dataset (Goyal et al. 2017), which is built based on the COCO (Lin et al. 2014) images.”
    From VL-BERT · §Experiment
  • Grid features for VQA2020 · cited 3×
    “Second, while our 1×\times1 RoIPool-based variant hurts the object detection performance (average precision lin2014microsoft on VG drops from 4.07 to 2.90), it helps VQA – boosting the accuracy by 0.73% (row 3 & 4) and a…”
    From Grid features for VQA · §Main Comparison: Regions vs. Grids
  • VL-BERT meta-analysis2020 · cited 2×, 2 in Method
    “COCO (Lin et al. 2014) or VQA (Antol et al. 2015), where the images are strongly-associated with crowdsourced captions or question--answer pairs.”
    From VL-BERT meta-analysis · §Experimental Setup
  • VinVL2021 · cited 5×, 1 in Method
    “We build our pre-training corpus based on three types of existing vision and VL datasets: (1) image captioning datasets with human-annotated captions as 𝒘\boldsymbol{w} and machine-generated 55 5 We use the same model t…”
    From VinVL · §Oscar+ Pre-training
  • VL-T52021 · cited 2×, 1 in Method
    “We aggregate pretraining data from MS COCO (Lin et al. 2014; Chen et al. 2015) and Visual Genome (VG; Krishna et al. 2016) images33 3 Existing vision-and-language transformers are trained with different datasets and comp…”
    From VL-T5 · §Pretraining
  • Conceptual 12M2021 · cited 4×, 2 in Method
    “Unlike in the standard image captioning setting, nocaps’s distributions of images during training (COCO Captions) and evaluation (Open Images) are different: the Open Images dataset [45, 42] covers one order of magnitude…”
    From Conceptual 12M · §Evaluating Vision-and-Language Pre-Training Data
  • DALL·E2021 · cited 1×, 1 in Method
    “Our preliminary experiments for models up to 1.21.2 billion parameters were carried out on Conceptual Captions, a dataset of 3.3 million text-image pairs that was developed as an extension to MS-COCO (Lin et al. 2014).”
    From DALL·E · §Method
  • CLIP2021 · cited 1×, 1 in Method
    “Existing work has mainly used three datasets, MS-COCO (Lin et al. 2014), Visual Genome (Krishna et al. 2017), and YFCC100M (Thomee et al. 2016).”
    From CLIP · §Approach
  • ALBEF2021 · cited 1×, 1 in Method
    “Following UNITER [2], we construct our pre-training data using two web datasets (Conceptual Captions [4], SBU Captions [5]) and two in-domain datasets (COCO [41] and Visual Genome [42]).”
    From ALBEF · §ALBEF Pre-training
  • VLMo2021 · cited 2×
    “Following previous work [3, 20], our pre-training data consists of four image captioning datasets: Conceptual Captions (CC) [39], SBU Captions [32], COCO [27] and Visual Genome (VG) [21] datasets.”
    From VLMo · §Experiments
  • METER2021 · cited 5×, 2 in Method
    “Vision-and-language (VL) tasks, such as visual question answering (VQA) antol2015vqa and image-text retrieval lin2014microsoft; plummer2015flickr30k, require an AI system to comprehend both the input image and text conte…”
    From METER · §Introduction
  • Swin V22021 · cited 2×
    “We conduct experiments on ImageNet-1K image classification (V1 and V2) deng2009imagenet; recht2019imagenet, COCO object detection lin2014coco, and ADE20K semantic segmentation zhou2018semantic.”
    From Swin V2 · §Experiments
  • ViTCAP2021 · cited 5×, 1 in Method
    “In particular, ViTCAP achieves 138.1138.1 CIDEr scores on COCO-caption Karpathy split lin2014microsoft, 108.6108.6 on Google-CC sharma2018conceptual, and 95.495.4 on nocaps agrawal2019nocaps datasets.”
    From ViTCAP · §Introduction
  • BLIP2022 · cited 1×, 1 in Method
    “Due to the prohibitive annotation cost, there exist a limited number of high-quality human-annotated image-text pairs {(Ih,Th)}\{(I_{h},T_{h})\} (e.g., COCO (Lin et al. 2014)).”
    From BLIP · §Method
  • Simple end-to-end captioning2022 · cited 4×, 1 in Method
    “For image captioning task, we use three different datasets, including MSCOCO Captions (Lin et al. 2015)44 4 https://cocodataset.org/#home, Flickr30k (Plummer et al. 2016)55 5 http://hockenmaier.cs.illinois.edu/Denotation…”
    From Simple end-to-end captioning · §. Experiment Setup
  • VL-BEiT2022 · cited 3×
    “For monomodal data, we use ImageNet-22K as the image data, English Wikipedia and BookCorpus [48] as the text data.”
    From VL-BEiT · §Experiments
  • BEiT-32022 · cited 6×, 1 in Method
    “For multimodal data, there are about 1515M images and 2121M image-text pairs collected from five public datasets: Conceptual 12M (CC12M) [11], Conceptual Captions (CC3M) [46], SBU Captions (SBU) [39], COCO [31] and Visua…”
    From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • EVA2022 · cited 4×
    “Using 29.6 million public accessible unlabeled images for pre-training, EVA sets new records on several representative vision benchmarks, such as image classification on ImageNet-1K deng2009imagenet (89.7% top-1 accuracy…”
    From EVA · §Introduction
  • BLIP-22023 · cited 1×, 1 in Method
    “We use the same pre-training dataset as BLIP with 129M images in total, including COCO (Lin et al. 2014), Visual Genome (Krishna et al. 2017), CC3M (Sharma et al. 2018), CC12M (Changpinyo et al. 2021), SBU (Ordonez et al…”
    From BLIP-2 · §Method
Abstract

We present a new dataset with the goal of advancing the state-of-the-art in object recognition by placing the question of object recognition in the context of the broader question of scene understanding. This is achieved by gathering images of complex everyday scenes containing common objects in their natural context. Objects are labeled using per-instance segmentations to aid in precise object localization. Our dataset contains photos of 91 objects types that would be easily recognizable by a 4 year old. With a total of 2.5 million labeled instances in 328k images, the creation of our dataset drew upon extensive crowd worker involvement via novel user interfaces for category detection, instance spotting and instance segmentation. We present a detailed statistical analysis of the dataset in comparison to PASCAL, ImageNet, and SUN. Finally, we provide baseline performance analysis for bounding box and segmentation detection results using a Deformable Parts Model.