Paper Lineage
Esc
MethodJun 2015arXiv 1506.01497cs.CV

Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

Shaoqing Ren, Kaiming He, Ross Girshick, Jian Sun

State-of-the-art object detection networks depend on region proposal algorithms to hypothesize object locations. Advances like SPPnet and Fast R-CNN have reduced the running time of these detection networks, exposing region proposal computation as a bottleneck.

From the abstract

Built on

1 paper · 0 verifiedSee as graph

Also cited · not yet reviewed (1)

  • ResNet2015 · cited 7×
    “In ILSVRC and COCO 2015 competitions, Faster R-CNN and RPN are the basis of several 1st-place entries He2015a in the tracks of ImageNet detection, ImageNet localization, COCO detection, and COCO segmentation.”
    From this paper · §I Introduction

Led to

  • ResNet2015 · cited 2×
    “We adopt Faster R-CNN Ren2015 as the detection method.”
    From ResNet · §Experiments
  • Visual Genome2016 · cited 2×
    “Much progress has been made in recent years towards this goal, including image classification Deng et al., 2009; Perronnin et al., 2010; Simonyan and Zisserman, 2014; Krizhevsky et al., 2012; Szegedy et al., 2014 and obj…”
    From Visual Genome · §Introduction
  • Goyal large-batch SGD2017 · cited 4×
    “Additionally, we show that the linear scaling rule and warmup generalize to more complex tasks including object detection and instance segmentation Girshick2015; Ren2015; He2017; Long2015, which we demonstrate via the re…”
    From Goyal large-batch SGD · §Introduction
  • JFT-300M (unreasonable effectiveness)2017 · cited 4×, 1 in Method
    “We use the Faster RCNN framework [33] for its state-of-the-art performance.”
    From JFT-300M (unreasonable effectiveness) · §Training and Evaluation Framework
  • NASNet2017 · cited 2×
    “In our experiments, the features learned by NASNets from ImageNet classification can be combined with the Faster-RCNN framework faster_rcnn to achieve state-of-the-art on COCO object detection task for both the largest a…”
    From NASNet · §Introduction
  • Bottom-Up Top-Down attention2017 · cited 3×, 1 in Method
    “Practically, we implement bottom-up attention using Faster R-CNN faster_rcnn, which represents a natural expression of a bottom-up attention mechanism.”
    From Bottom-Up Top-Down attention · §Introduction
  • Neural Baby Talk2018 · cited 4×, 3 in Method
    “Given an image 𝑰\bm{I}, and the corresponding caption 𝒚\bm{y}, the candidate grounding regions are obtained by using a pre-trained Faster-RCNN network ren2015faster.”
    From Neural Baby Talk · §Method
  • “Note that both Query-Adaptive RCNN [10] and our off-the-shelf object detector [2] are based on Faster RCNN [25] and pre-trained on Visual Genome [17].”
    From Bilinear Attention Networks · §Flickr30k entities results and discussions
  • nocaps2018 · cited 3×, 2 in Method
    “Hence, we use a Faster-RCNN [36] model pre-trained using Open Images V4 [20] (referred as OI detector henceforth), to obtain candidate region proposals as described in Section 4 of the main paper.”
    From nocaps · §Additional Implementation Details for Baseline Models
  • CPC v22019 · cited 2×, 1 in Method
    “As such 𝔻l{\mathbb{D}}_{l} is the entire PASCAL VOC 2007 dataset (comprised of 5011 labeled images); hψh_{\psi} and ℒSupL_{\textrm{Sup}} are the Faster-RCNN architecture and loss (Ren et al. 2015).”
    From CPC v2 · §Experimental Setup
  • LXMERT2019 · cited 2×, 2 in Method
    “For these reasons, we take detected labels output by Faster R-CNN Ren et al. 2015.”
    From LXMERT · §Pre-Training Strategies
  • Decoupled box proposals captioning2019 · cited 1×, 1 in Method
    “In this paper, we select Faster R-CNN Ren et al. 2015b, a widely-used object detector in image captioning and VQA.”
    From Decoupled box proposals captioning · §Features and Experimental Setup
  • MoCo2019 · cited 2×
    “The detector is Faster R-CNN Ren2015 with a backbone of R50-dilated-C5 or R50-C4 He2017 (details in appendix), with BN tuned, implemented in Wu2019.”
    From MoCo · §Experiments
  • Grid features for VQA2020 · cited 6×
    “Unlike normal attention that uses ‘top-down’ linguistic inputs to focus on specific parts of the visual input, bottom-up attention uses pre-trained object detectors ren2015faster to identify salient regions based solely…”
    From Grid features for VQA · §Introduction
  • VL-BERT meta-analysis2020 · cited 1×, 1 in Method
    “V&L BERTs typically extract image features using a Faster R-CNN (Ren et al. 2015) trained on the Visual Genome dataset (VG; Krishna et al. 2017), either with a ResNet-101101 (He et al. 2016) or a ResNeXT-152152 backbone…”
    From VL-BERT meta-analysis · §Experimental Setup
  • VL-T52021 · cited 1×, 1 in Method
    “We represent an input image vv with n=36n{=}36 object regions from a Faster R-CNN (Ren et al. 2015) trained on Visual Genome (Krishna et al. 2016) for object and attribute classification (Anderson et al. 2018).”
    From VL-T5 · §Model
  • ViLT2021 · cited 1×, 1 in Method
    “They are obtained from an off-the-shelf object detector like Faster R-CNN (Ren et al. 2016).”
    From ViLT · §Background
  • Conceptual 12M2021 · cited 1×, 1 in Method
    “We train a Faster-RCNN [68] on Visual Genome [43], with a ResNet101 [29] backbone trained on JFT [30] and fine-tuned on ImageNet [69].”
    From Conceptual 12M · §Evaluating Vision-and-Language Pre-Training Data
  • Dynamic Head2021 · cited 4×, 1 in Method
    “Two-stage detectors utilize region proposal and ROI-pooling [23] layers to extract intermediate representations from feature pyramid of a backbone network.”
    From Dynamic Head · §Our Approach
Abstract

State-of-the-art object detection networks depend on region proposal algorithms to hypothesize object locations. Advances like SPPnet and Fast R-CNN have reduced the running time of these detection networks, exposing region proposal computation as a bottleneck. In this work, we introduce a Region Proposal Network (RPN) that shares full-image convolutional features with the detection network, thus enabling nearly cost-free region proposals. An RPN is a fully convolutional network that simultaneously predicts object bounds and objectness scores at each position. The RPN is trained end-to-end to generate high-quality region proposals, which are used by Fast R-CNN for detection. We further merge RPN and Fast R-CNN into a single network by sharing their convolutional features---using the recently popular terminology of neural networks with 'attention' mechanisms, the RPN component tells the unified network where to look. For the very deep VGG-16 model, our detection system has a frame rate of 5fps (including all steps) on a GPU, while achieving state-of-the-art object detection accuracy on PASCAL VOC 2007, 2012, and MS COCO datasets with only 300 proposals per image. In ILSVRC and COCO 2015 competitions, Faster R-CNN and RPN are the foundations of the 1st-place winning entries in several tracks. Code has been made publicly available.