The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale
2M images with unified annotations for image classification, object detection and visual relationship detection. The images have a Creative Commons Attribution license that allows to share and adapt the material, and they have been collected from Flickr without a predefined list of class names or tags, leading to natural class statistics and avoiding an initial design bias.
Also cited · not yet reviewed (1)
- Box Attention2018 · cited 4×, 3 in Method“Recently, high-performing models based on deep convolutional neural networks are dominating the field Gupta and Malik 2015; Gkioxari et al. 2018; Gao et al. 2018; Kolesnikov et al. 2018.”From this paper · §Performance of baseline models
Led to
- Box Attention2018 · cited 2דWe evaluate our approach on the V-COCO [10], Visual Relationships [20] and Open Images [16] datasets in Section 4.”From Box Attention · §Introduction
- Localized Narratives2019 · cited 2דWe collected Localized Narratives at scale: we annotated the whole COCO [35] (123123k images), ADE20K [69] (2020k) and Flickr30k [66] (3232k) datasets, as well as 671671k images of Open Images [33].”From Localized Narratives · §Introduction
- Oscar2020 · cited 2דTo study the impact of different object tag sets in pre-trained models, we pre-train two variants: OscarVG^{VG} and OscarOI^{OI} utilizes object tags produced by the object detector trained on the visual genome (VG) data…”From Oscar · §Experimental Results & Analysis
- VinVL2021 · cited 2דCompared to the OD model of [2], the new model is better-designed for VL tasks, and is bigger and trained on much larger amounts of data, combining multiple public object detection datasets, including COCO [25], OpenImag…”From VinVL · §Introduction
- Conceptual 12M2021 · cited 2×, 2 in Method“Unlike in the standard image captioning setting, nocaps’s distributions of images during training (COCO Captions) and evaluation (Open Images) are different: the Open Images dataset [45, 42] covers one order of magnitude…”From Conceptual 12M · §Evaluating Vision-and-Language Pre-Training Data
Abstract
We present Open Images V4, a dataset of 9.2M images with unified annotations for image classification, object detection and visual relationship detection. The images have a Creative Commons Attribution license that allows to share and adapt the material, and they have been collected from Flickr without a predefined list of class names or tags, leading to natural class statistics and avoiding an initial design bias. Open Images V4 offers large scale across several dimensions: 30.1M image-level labels for 19.8k concepts, 15.4M bounding boxes for 600 object classes, and 375k visual relationship annotations involving 57 classes. For object detection in particular, we provide 15x more bounding boxes than the next largest datasets (15.4M boxes on 1.9M images). The images often show complex scenes with several objects (8 annotated objects per image on average). We annotated visual relationships between them, which support visual relationship detection, an emerging task that requires structured reasoning. We provide in-depth comprehensive statistics about the dataset, we validate the quality of the annotations, we study how the performance of several modern models evolves with increasing amounts of training data, and we demonstrate two applications made possible by having unified annotations of multiple types coexisting in the same images. We hope that the scale, quality, and variety of Open Images V4 will foster further research and innovation even beyond the areas of image classification, object detection, and visual relationship detection.