LVIS: A Dataset for Large Vocabulary Instance Segmentation
Progress on object detection is enabled by datasets that focus the research community's attention on open challenges. This process led us from simple images to complex scenes and from bounding boxes to segmentation masks.
Also cited · not yet reviewed (1)
- MS COCO2014 · cited 2דQuality is important for future research because relatively coarse masks, such as those in the COCO dataset Lin2014, limit the ability to differentiate algorithm-predicted mask quality beyond a certain, coarse point.”From this paper · §Introduction
Led to
- Florence2021 · cited 1×, 1 in Method“We merge several well-known object detection datasets, including COCO (Lin et al. 2015), LVIS (Gupta et al. 2019), OpenImages (Krasin et al. 2016), Object365 (Shao et al. 2019).”From Florence · §Approach
- EVA2022 · cited 6דUsing 29.6 million public accessible unlabeled images for pre-training, EVA sets new records on several representative vision benchmarks, such as image classification on ImageNet-1K deng2009imagenet (89.7% top-1 accuracy…”From EVA · §Introduction
Abstract
Progress on object detection is enabled by datasets that focus the research community's attention on open challenges. This process led us from simple images to complex scenes and from bounding boxes to segmentation masks. In this work, we introduce LVIS (pronounced `el-vis'): a new dataset for Large Vocabulary Instance Segmentation. We plan to collect ~2 million high-quality instance segmentation masks for over 1000 entry-level object categories in 164k images. Due to the Zipfian distribution of categories in natural images, LVIS naturally has a long tail of categories with few training samples. Given that state-of-the-art deep learning methods for object detection perform poorly in the low-sample regime, we believe that our dataset poses an important and exciting new scientific challenge. LVIS is available at http://www.lvisdataset.org.