nocaps: novel object captioning at scale
Image captioning models have achieved impressive results on datasets containing limited visual concepts and large amounts of paired image-caption training data. However, if these models are to ever function in the wild, a much larger variety of visual concepts must be learned, ideally from less supervision.
Also cited · not yet reviewed (7)
- Neural Baby Talk2018 · cited 10×, 3 in Method“To establish the state-of-the-art on our challenging benchmark, we evaluate two of the best performing existing approaches [27, 2] and report their performance based on well-established evaluation metrics – CIDEr [41] an…”From this paper · §Introduction
- Constrained beam search captioning2016 · cited 10×, 2 in Method“To establish the state-of-the-art on our challenging benchmark, we evaluate two of the best performing existing approaches [27, 2] and report their performance based on well-established evaluation metrics – CIDEr [41] an…”From this paper · §Introduction
- Faster R-CNN2015 · cited 3×, 2 in Method“Hence, we use a Faster-RCNN [36] model pre-trained using Open Images V4 [20] (referred as OI detector henceforth), to obtain candidate region proposals as described in Section 4 of the main paper.”From this paper · §Additional Implementation Details for Baseline Models
- Bottom-Up Top-Down attention2017 · cited 5×, 1 in Method“To provide an initial measure of the state-of-the-art on nocaps, we extend and present results for two contemporary approaches to novel object captioning – Neural Baby Talk (NBT) [27] and Constrained Beam Search (CBS) [2…”From this paper · §Experiments
Show 3 more
- ELMo2018 · cited 2×, 1 in Method“When using ELMo [34], we use a dynamic representation of wcw_{c}, h¯t1\bar{h}_{t}^{1} and h¯t2\bar{h}_{t}^{2} as the input word embedding wELMotw_{ELMo}^{t} for our caption model. wcw_{c} is the character embedding of…”From this paper · §Additional Implementation Details for Baseline Models
- COCO Captions2015 · cited 8דTest images contain previously unseen, or ‘novel’ objects that are drawn from a target distribution (in this case, Open Images [20]) that differs from the source/training distribution (COCO [6]).”From this paper · §Related Work
- CIDEr2014 · cited 4דTo establish the state-of-the-art on our challenging benchmark, we evaluate two of the best performing existing approaches [27, 2] and report their performance based on well-established evaluation metrics – CIDEr [41] an…”From this paper · §Introduction
Led to
- Meshed-Memory Transformer2019 · cited 2דThen, we assess the captioning of novel objects by testing on the recently proposed nocaps dataset agrawal2019nocaps.”From Meshed-Memory Transformer · §Experiments
- VinVL2021 · cited 2דTo validate the effectiveness of the new OD model, we pre-train a Transformer-based cross-modal fusion model Oscar+ [21] on a public dataset consisting of 8.858.85 million text-image pairs, where the visual representatio…”From VinVL · §Introduction
- Conceptual 12M2021 · cited 4×, 1 in Method“Table 4 compares our best model (ic pre-trained on CC3M+CC12M) to existing state-of-the-art results on nocaps, and show that ours achieves state-of-the-art performance on CIDEr, outperforming a concurrent work [32] that…”From Conceptual 12M · §Experimental Results
- ViTCAP2021 · cited 4דIn particular, ViTCAP achieves 138.1138.1 CIDEr scores on COCO-caption Karpathy split lin2014microsoft, 108.6108.6 on Google-CC sharma2018conceptual, and 95.495.4 on nocaps agrawal2019nocaps datasets.”From ViTCAP · §Introduction
- Simple end-to-end captioning2022 · cited 2×, 1 in Method“For image captioning task, we use three different datasets, including MSCOCO Captions (Lin et al. 2015)44 4 https://cocodataset.org/#home, Flickr30k (Plummer et al. 2016)55 5 http://hockenmaier.cs.illinois.edu/Denotation…”From Simple end-to-end captioning · §. Experiment Setup
Abstract
Image captioning models have achieved impressive results on datasets containing limited visual concepts and large amounts of paired image-caption training data. However, if these models are to ever function in the wild, a much larger variety of visual concepts must be learned, ideally from less supervision. To encourage the development of image captioning models that can learn visual concepts from alternative data sources, such as object detection datasets, we present the first large-scale benchmark for this task. Dubbed 'nocaps', for novel object captioning at scale, our benchmark consists of 166,100 human-generated captions describing 15,100 images from the OpenImages validation and test sets. The associated training data consists of COCO image-caption pairs, plus OpenImages image-level labels and object bounding boxes. Since OpenImages contains many more classes than COCO, nearly 400 object classes seen in test images have no or very few associated training captions (hence, nocaps). We extend existing novel object captioning models to establish strong baselines for this benchmark and provide analysis to guide future work on this task.