Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
Rectified activation units (rectifiers) are essential for state-of-the-art neural networks. In this work, we study rectifier neural networks for image classification from two aspects.
Also cited · not yet reviewed (4)
- VGG2014 · cited 21×, 14 in Method“With fixed standard deviations (e.g., 0.01 in Krizhevsky2012), very deep models (e.g., >>8 conv layers) have difficulties to converge, as reported by the VGG team Simonyan2014 and also observed in our experiments.”From this paper · §Approach
- OverFeat2013 · cited 5×, 2 in Method“We further improve this strategy using the dense sliding window method in Sermanet2014; Simonyan2014.”From this paper · §Implementation Details
- GoogLeNet (Inception)2014 · cited 7×, 1 in Method“In Szegedy2014; Lee2014, auxiliary classifiers are added to intermediate layers to help with convergence.”From this paper · §Approach
- Return of the Devil (CNN details)2014 · cited 2×, 1 in Method“Our training algorithm mostly follows Krizhevsky2012; Howard2013; Chatfield2014; He2014; Simonyan2014.”From this paper · §Implementation Details
Led to
- Inception v32015 · cited 2דStill our solution uses much less computation than the best published results based on denser networks: our model outperforms the results of He et al he2015delving – cutting the top-55 (top-11) error by 25%25\% (14%14\%)…”From Inception v3 · §Conclusions
- ResNet2015 · cited 5דThis problem, however, has been largely addressed by normalized initialization LeCun1998; Glorot2010; Saxe2013; He2015 and intermediate normalization layers Ioffe2015, which enable networks with tens of layers to start c…”From ResNet · §Introduction
- ResNeXt2016 · cited 1×, 1 in Method“We adopt the weight initialization of He2015.”From ResNeXt · §Implementation details
- Instagram hashtag pre-training2018 · cited 1×, 1 in Method“We use a warm-up from 0.1 up to 0.1/256×80640.1/256\times 8064, where 0.1 and 256 are canonical learning rate and minibatch sizes [28].”From Instagram hashtag pre-training · §Scaling up Supervised Pretraining
- ImageNetV22019 · cited 2דState-of-the-art models now surpass human-level accuracy by some measure [20, 48].”From ImageNetV2 · §Summary of Our Experiments
Abstract
Rectified activation units (rectifiers) are essential for state-of-the-art neural networks. In this work, we study rectifier neural networks for image classification from two aspects. First, we propose a Parametric Rectified Linear Unit (PReLU) that generalizes the traditional rectified unit. PReLU improves model fitting with nearly zero extra computational cost and little overfitting risk. Second, we derive a robust initialization method that particularly considers the rectifier nonlinearities. This method enables us to train extremely deep rectified models directly from scratch and to investigate deeper or wider network architectures. Based on our PReLU networks (PReLU-nets), we achieve 4.94% top-5 test error on the ImageNet 2012 classification dataset. This is a 26% relative improvement over the ILSVRC 2014 winner (GoogLeNet, 6.66%). To our knowledge, our result is the first to surpass human-level performance (5.1%, Russakovsky et al.) on this visual recognition challenge.