Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Training Deep Neural Networks is complicated by the fact that the distribution of each layer's inputs changes during training, as the parameters of the previous layers change. This slows down the training by requiring lower learning rates and careful parameter initialization, and makes it notoriously hard to train models with saturating nonlinearities.
Also cited · not yet reviewed (1)
- GoogLeNet (Inception)2014 · cited 4דThe details of ensemble and multicrop inference are similar to Szegedy et al. 2014.”From this paper · §Experiments
Led to
- Inception v32015 · cited 4דThis is achieved with relatively modest (2.5×2.5\times) increase in computational cost compared to the network described in Ioffe et al ioffe2015batch.”From Inception v3 · §Conclusions
- ResNet2015 · cited 8דThis problem, however, has been largely addressed by normalized initialization LeCun1998; Glorot2010; Saxe2013; He2015 and intermediate normalization layers Ioffe2015, which enable networks with tens of layers to start c…”From ResNet · §Introduction
- ResNeXt2016 · cited 3×, 1 in Method“Unlike VGG-nets, the family of Inception models Szegedy2015; Ioffe2015; Szegedy2016a; Szegedy2016 have demonstrated that carefully designed topologies are able to achieve compelling accuracy with low theoretical complexi…”From ResNeXt · §Introduction
- MobileNets2017 · cited 3×, 1 in Method“MobileNets are built primarily from depthwise separable convolutions initially introduced in sifre2014rigid and subsequently used in Inception models ioffe2015batch to reduce the computation in the first few layers.”From MobileNets · §Prior Work
- Goyal large-batch SGD2017 · cited 3דWe use a weight decay λ\lambda of 0.0001 and following He2016 we do not apply weight decay on the learnable BN coefficients (namely, γ\gamma and β\beta in Ioffe2015).”From Goyal large-batch SGD · §Main Results and Analysis
- JFT-300M (unreasonable effectiveness)2017 · cited 1×, 1 in Method“We set weight decay to 10−410^{-4} and use batch normalization [20] after all the convolutional layers.”From JFT-300M (unreasonable effectiveness) · §Training and Evaluation Framework
- NASNet2017 · cited 3דThanks to this property of the cells, we can generate a family of models that achieve accuracies superior to all human-invented models at equivalent or smaller computational budgets szegedy2016rethinking; BatchNorm.”From NASNet · §Introduction
- Instagram hashtag pre-training2018 · cited 1×, 1 in Method“Each GPU processes 24 images at a time and batch normalization (BN) [27] statistics are computed on these 24 image sets.”From Instagram hashtag pre-training · §Scaling up Supervised Pretraining
- Billion-scale semi-supervised2019 · cited 2דEach GPU processes 2424 images at a time and apply batch normalization [22] to all convolutional layers on each GPU.”From Billion-scale semi-supervised · §Image classification: experiments & analysis
- EfficientNet2019 · cited 1×, 1 in Method“Although several techniques, such as skip connections (He et al. 2016) and batch normalization (Ioffe & Szegedy 2015), alleviate the training problem, the accuracy gain of very deep network diminishes: for example, ResNe…”From EfficientNet · §Compound Model Scaling
- MoCo2019 · cited 1×, 1 in Method“Our encoders fqf_{\textrm{q}} and fkf_{\textrm{k}} both have Batch Normalization (BN) Ioffe2015 as in the standard ResNet He2016.”From MoCo · §Method
- SimCLR2020 · cited 1×, 1 in Method“Standard ResNets use batch normalization (Ioffe & Szegedy 2015).”From SimCLR · §Method
- BYOL2020 · cited 1×, 1 in Method“This MLP consists in a linear layer with output size 40964096 followed by batch normalization [68], rectified linear units (ReLU) [69], and a final linear layer with output dimension 256256.”From BYOL · §Method
Abstract
Training Deep Neural Networks is complicated by the fact that the distribution of each layer's inputs changes during training, as the parameters of the previous layers change. This slows down the training by requiring lower learning rates and careful parameter initialization, and makes it notoriously hard to train models with saturating nonlinearities. We refer to this phenomenon as internal covariate shift, and address the problem by normalizing layer inputs. Our method draws its strength from making normalization a part of the model architecture and performing the normalization for each training mini-batch. Batch Normalization allows us to use much higher learning rates and be less careful about initialization. It also acts as a regularizer, in some cases eliminating the need for Dropout. Applied to a state-of-the-art image classification model, Batch Normalization achieves the same accuracy with 14 times fewer training steps, and beats the original model by a significant margin. Using an ensemble of batch-normalized networks, we improve upon the best published result on ImageNet classification: reaching 4.9% top-5 validation error (and 4.8% test error), exceeding the accuracy of human raters.