Decoupled Box Proposal and Featurization with Ultrafine-Grained Semantic Labels Improve Image Captioning and Visual Question Answering
Object detection plays an important role in current solutions to vision and language tasks like image captioning and visual question answering. However, popular models like Faster R-CNN rely on a costly process of annotating ground-truths for both the bounding boxes and their corresponding semantic labels, making it less amenable as a primitive task for transfer learning.
Also cited · not yet reviewed (6)
- Bottom-Up Top-Down attention2017 · cited 6×, 2 in Method“We reimplement the Faster R-CNN model, training it to predict both 1,600 object and 400 attribute labels in Visual Genome Krishna et al. 2017, following the standard setting from Anderson et al. 2018.”From this paper · §Features and Experimental Setup
- Visual Genome2016 · cited 2×, 1 in Method“We reimplement the Faster R-CNN model, training it to predict both 1,600 object and 400 attribute labels in Visual Genome Krishna et al. 2017, following the standard setting from Anderson et al. 2018.”From this paper · §Features and Experimental Setup
- VQA v22016 · cited 2×, 1 in Method“Other VQA datasets, including VQA1.0 Antol et al. 2015, VQA2.0 Goyal et al. 2017, Visual7W Zhu et al. 2016, COCOQA Ren et al. 2015a, and GQA Hudson and Manning 2019 are completely or partly based on MSCOCO or Visual Geno…”From this paper · §Visual Question Answering
- Faster R-CNN2015 · cited 1×, 1 in Method“In this paper, we select Faster R-CNN Ren et al. 2015b, a widely-used object detector in image captioning and VQA.”From this paper · §Features and Experimental Setup
Show 2 more
- ResNet2015 · cited 1×, 1 in Method“ResNet-101 He et al. 2016 pre-trained on ImageNet Russakovsky et al. 2015 is used as the core featurization network11 1 See further details in the supplementary material..”From this paper · §Features and Experimental Setup
- VQA2015 · cited 2דOther VQA datasets, including VQA1.0 Antol et al. 2015, VQA2.0 Goyal et al. 2017, Visual7W Zhu et al. 2016, COCOQA Ren et al. 2015a, and GQA Hudson and Manning 2019 are completely or partly based on MSCOCO or Visual Geno…”From this paper · §Visual Question Answering
Led to
- Conceptual 12M2021 · cited 6×, 4 in Method“We select top-16 box proposals and featurize each of them with Graph-RISE, similar to [19].”From Conceptual 12M · §Evaluating Vision-and-Language Pre-Training Data
Abstract
Object detection plays an important role in current solutions to vision and language tasks like image captioning and visual question answering. However, popular models like Faster R-CNN rely on a costly process of annotating ground-truths for both the bounding boxes and their corresponding semantic labels, making it less amenable as a primitive task for transfer learning. In this paper, we examine the effect of decoupling box proposal and featurization for down-stream tasks. The key insight is that this allows us to leverage a large amount of labeled annotations that were previously unavailable for standard object detection benchmarks. Empirically, we demonstrate that this leads to effective transfer learning and improved image captioning and visual question answering models, as measured on publicly available benchmarks.