Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
State-of-the-art object detection networks depend on region proposal algorithms to hypothesize object locations. Advances like SPPnet and Fast R-CNN have reduced the running time of these detection networks, exposing region proposal computation as a bottleneck.
Also cited · not yet reviewed (1)
- ResNet2015 · cited 7דIn ILSVRC and COCO 2015 competitions, Faster R-CNN and RPN are the basis of several 1st-place entries He2015a in the tracks of ImageNet detection, ImageNet localization, COCO detection, and COCO segmentation.”From this paper · §I Introduction
Led to
- ResNet2015 · cited 2דWe adopt Faster R-CNN Ren2015 as the detection method.”From ResNet · §Experiments
- Visual Genome2016 · cited 2דMuch progress has been made in recent years towards this goal, including image classification Deng et al., 2009; Perronnin et al., 2010; Simonyan and Zisserman, 2014; Krizhevsky et al., 2012; Szegedy et al., 2014 and obj…”From Visual Genome · §Introduction
- Goyal large-batch SGD2017 · cited 4דAdditionally, we show that the linear scaling rule and warmup generalize to more complex tasks including object detection and instance segmentation Girshick2015; Ren2015; He2017; Long2015, which we demonstrate via the re…”From Goyal large-batch SGD · §Introduction
- JFT-300M (unreasonable effectiveness)2017 · cited 4×, 1 in Method“We use the Faster RCNN framework [33] for its state-of-the-art performance.”From JFT-300M (unreasonable effectiveness) · §Training and Evaluation Framework
- NASNet2017 · cited 2דIn our experiments, the features learned by NASNets from ImageNet classification can be combined with the Faster-RCNN framework faster_rcnn to achieve state-of-the-art on COCO object detection task for both the largest a…”From NASNet · §Introduction
- Bottom-Up Top-Down attention2017 · cited 3×, 1 in Method“Practically, we implement bottom-up attention using Faster R-CNN faster_rcnn, which represents a natural expression of a bottom-up attention mechanism.”From Bottom-Up Top-Down attention · §Introduction
- Neural Baby Talk2018 · cited 4×, 3 in Method“Given an image 𝑰\bm{I}, and the corresponding caption 𝒚\bm{y}, the candidate grounding regions are obtained by using a pre-trained Faster-RCNN network ren2015faster.”From Neural Baby Talk · §Method
- Bilinear Attention Networks2018 · cited 2דNote that both Query-Adaptive RCNN [10] and our off-the-shelf object detector [2] are based on Faster RCNN [25] and pre-trained on Visual Genome [17].”From Bilinear Attention Networks · §Flickr30k entities results and discussions
- nocaps2018 · cited 3×, 2 in Method“Hence, we use a Faster-RCNN [36] model pre-trained using Open Images V4 [20] (referred as OI detector henceforth), to obtain candidate region proposals as described in Section 4 of the main paper.”From nocaps · §Additional Implementation Details for Baseline Models
- CPC v22019 · cited 2×, 1 in Method“As such 𝔻l{\mathbb{D}}_{l} is the entire PASCAL VOC 2007 dataset (comprised of 5011 labeled images); hψh_{\psi} and ℒSupL_{\textrm{Sup}} are the Faster-RCNN architecture and loss (Ren et al. 2015).”From CPC v2 · §Experimental Setup
- LXMERT2019 · cited 2×, 2 in Method“For these reasons, we take detected labels output by Faster R-CNN Ren et al. 2015.”From LXMERT · §Pre-Training Strategies
- Decoupled box proposals captioning2019 · cited 1×, 1 in Method“In this paper, we select Faster R-CNN Ren et al. 2015b, a widely-used object detector in image captioning and VQA.”From Decoupled box proposals captioning · §Features and Experimental Setup
- MoCo2019 · cited 2דThe detector is Faster R-CNN Ren2015 with a backbone of R50-dilated-C5 or R50-C4 He2017 (details in appendix), with BN tuned, implemented in Wu2019.”From MoCo · §Experiments
- Grid features for VQA2020 · cited 6דUnlike normal attention that uses ‘top-down’ linguistic inputs to focus on specific parts of the visual input, bottom-up attention uses pre-trained object detectors ren2015faster to identify salient regions based solely…”From Grid features for VQA · §Introduction
- VL-BERT meta-analysis2020 · cited 1×, 1 in Method“V&L BERTs typically extract image features using a Faster R-CNN (Ren et al. 2015) trained on the Visual Genome dataset (VG; Krishna et al. 2017), either with a ResNet-101101 (He et al. 2016) or a ResNeXT-152152 backbone…”From VL-BERT meta-analysis · §Experimental Setup
- VL-T52021 · cited 1×, 1 in Method“We represent an input image vv with n=36n{=}36 object regions from a Faster R-CNN (Ren et al. 2015) trained on Visual Genome (Krishna et al. 2016) for object and attribute classification (Anderson et al. 2018).”From VL-T5 · §Model
- ViLT2021 · cited 1×, 1 in Method“They are obtained from an off-the-shelf object detector like Faster R-CNN (Ren et al. 2016).”From ViLT · §Background
- Conceptual 12M2021 · cited 1×, 1 in Method“We train a Faster-RCNN [68] on Visual Genome [43], with a ResNet101 [29] backbone trained on JFT [30] and fine-tuned on ImageNet [69].”From Conceptual 12M · §Evaluating Vision-and-Language Pre-Training Data
- Dynamic Head2021 · cited 4×, 1 in Method“Two-stage detectors utilize region proposal and ROI-pooling [23] layers to extract intermediate representations from feature pyramid of a backbone network.”From Dynamic Head · §Our Approach
Abstract
State-of-the-art object detection networks depend on region proposal algorithms to hypothesize object locations. Advances like SPPnet and Fast R-CNN have reduced the running time of these detection networks, exposing region proposal computation as a bottleneck. In this work, we introduce a Region Proposal Network (RPN) that shares full-image convolutional features with the detection network, thus enabling nearly cost-free region proposals. An RPN is a fully convolutional network that simultaneously predicts object bounds and objectness scores at each position. The RPN is trained end-to-end to generate high-quality region proposals, which are used by Fast R-CNN for detection. We further merge RPN and Fast R-CNN into a single network by sharing their convolutional features---using the recently popular terminology of neural networks with 'attention' mechanisms, the RPN component tells the unified network where to look. For the very deep VGG-16 model, our detection system has a frame rate of 5fps (including all steps) on a GPU, while achieving state-of-the-art object detection accuracy on PASCAL VOC 2007, 2012, and MS COCO datasets with only 300 proposals per image. In ILSVRC and COCO 2015 competitions, Faster R-CNN and RPN are the foundations of the 1st-place winning entries in several tracks. Code has been made publicly available.