OverFeat: Integrated Recognition, Localization and Detection using Convolutional Networks
We present an integrated framework for using Convolutional Networks for classification, localization and detection. We show how a multiscale and sliding window approach can be efficiently implemented within a ConvNet.
OverFeat has no earlier papers in this dataset.
Led to
- CNN features off-the-shelf2014 · cited 3×, 2 in Method“In this work we use the publicly available trained CNN called OverFeat Sermanet13.”From CNN features off-the-shelf · §Background and Outline
- VGG2014 · cited 9×, 2 in Method“Convolutional networks (ConvNets) have recently enjoyed a great success in large-scale image and video recognition (Krizhevsky et al. 2012; Zeiler & Fergus 2013; Sermanet et al. 2014; Simonyan & Zisserman 2014) which has…”From VGG · §Introduction
- GoogLeNet (Inception)2014 · cited 3דFor larger datasets such as Imagenet, the recent trend has been to increase the number of layers [12] and layer size [21, 14], while using dropout [7] to address the problem of overfitting.”From GoogLeNet (Inception) · §Related Work
- PReLU / He init2015 · cited 5×, 2 in Method“We further improve this strategy using the dense sliding window method in Sermanet2014; Simonyan2014.”From PReLU / He init · §Implementation Details
- Visual Genome2016 · cited 2דMuch progress has been made in recent years towards this goal, including image classification Deng et al., 2009; Perronnin et al., 2010; Simonyan and Zisserman, 2014; Krizhevsky et al., 2012; Szegedy et al., 2014 and obj…”From Visual Genome · §Introduction
- Goyal large-batch SGD2017 · cited 2דWe are in an unprecedented era in AI research history in which the increasing data and model scale is rapidly improving accuracy in computer vision Krizhevsky2012; Zeiler2014; Sermanet2014; Simonyan2015; Szegedy2015; He2…”From Goyal large-batch SGD · §Introduction
Abstract
We present an integrated framework for using Convolutional Networks for classification, localization and detection. We show how a multiscale and sliding window approach can be efficiently implemented within a ConvNet. We also introduce a novel deep learning approach to localization by learning to predict object boundaries. Bounding boxes are then accumulated rather than suppressed in order to increase detection confidence. We show that different tasks can be learned simultaneously using a single shared network. This integrated framework is the winner of the localization task of the ImageNet Large Scale Visual Recognition Challenge 2013 (ILSVRC2013) and obtained very competitive results for the detection and classifications tasks. In post-competition work, we establish a new state of the art for the detection task. Finally, we release a feature extractor from our best model called OverFeat.