Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Deep learning thrives with large neural networks and large datasets. However, larger networks and larger datasets result in longer training times that impede research and development progress.
Also cited · not yet reviewed (7)
- ResNet2015 · cited 20דOf equal importance, in a research domain, we have found it to simplify migrating algorithms from a single-GPU to a multi-GPU implementation without requiring hyper-parameter search, e.g. in our experience migrating Fast…”From this paper · §Introduction
- Faster R-CNN2015 · cited 4דAdditionally, we show that the linear scaling rule and warmup generalize to more complex tasks including object detection and instance segmentation Girshick2015; Ren2015; He2017; Long2015, which we demonstrate via the re…”From this paper · §Introduction
- GoogLeNet (Inception)2014 · cited 3דWe use scale and aspect ratio data augmentation Szegedy2015 as in Gross2016.”From this paper · §Main Results and Analysis
- BatchNorm2015 · cited 3דWe use a weight decay λ\lambda of 0.0001 and following He2016 we do not apply weight decay on the learnable BN coefficients (namely, γ\gamma and β\beta in Ioffe2015).”From this paper · §Main Results and Analysis
Show 3 more
- OverFeat2013 · cited 2דWe are in an unprecedented era in AI research history in which the increasing data and model scale is rapidly improving accuracy in computer vision Krizhevsky2012; Zeiler2014; Sermanet2014; Simonyan2015; Szegedy2015; He2…”From this paper · §Introduction
- VGG2014 · cited 2דWe are in an unprecedented era in AI research history in which the increasing data and model scale is rapidly improving accuracy in computer vision Krizhevsky2012; Zeiler2014; Sermanet2014; Simonyan2015; Szegedy2015; He2…”From this paper · §Introduction
- ResNeXt2016 · cited 2דTo investigate this question, we adopt the ImageNet-5k dataset suggested by Xie et al. xie2017 that extends ImageNet-1k to 6.8 million images (roughly 5×\times larger) by adding 4k additional categories from ImageNet-22k…”From this paper · §Main Results and Analysis
Led to
- Instagram hashtag pre-training2018 · cited 4×, 4 in Method“To make training at this scale practical, we adopt a distributed synchronous implementation of stochastic gradient descent with large (8k image) minibatches, following Goyal et al. [19].”From Instagram hashtag pre-training · §Scaling up Supervised Pretraining
- MnasNet2018 · cited 1×, 1 in Method“Following imagenet1hour17, learning rate is increased from 0 to 0.256 in the first 5 epochs, and then decayed by 0.97 every 2.4 epochs.”From MnasNet · §Experimental Setup
- Billion-scale semi-supervised2019 · cited 2דWe set the learning rate following the linear scaling procedure proposed in [13] with a warm-up and overall minibatch size of 64×24=153664\times 24=1536.”From Billion-scale semi-supervised · §Image classification: experiments & analysis
- Megatron-LM2019 · cited 1×, 1 in Method“Further research (Goyal et al. 2017; You et al. 2017; You et al. 2019) has developed techniques to mitigate these effects and drive down the training time of large neural networks.”From Megatron-LM · §Background and Challenges
- MoCo2019 · cited 4×, 1 in Method“It is also challenged by large mini-batch optimization Goyal2017.”From MoCo · §Method
- SimCLR2020 · cited 2×, 1 in Method“In contrast to supervised learning (Goyal et al. 2017), in contrastive learning, larger batch sizes provide more negative examples, facilitating convergence (i.e. taking fewer epochs and steps for a given accuracy).”From SimCLR · §Loss Functions and Batch Size
- BYOL2020 · cited 1×, 1 in Method“We set the base learning rate to 0.2,0.2, scaled linearly [72] with the batch size (LearningRate=0.2×BatchSize/256LearningRate=0.2\timesBatchSize/256).”From BYOL · §Method
Abstract
Deep learning thrives with large neural networks and large datasets. However, larger networks and larger datasets result in longer training times that impede research and development progress. Distributed synchronous SGD offers a potential solution to this problem by dividing SGD minibatches over a pool of parallel workers. Yet to make this scheme efficient, the per-worker workload must be large, which implies nontrivial growth in the SGD minibatch size. In this paper, we empirically show that on the ImageNet dataset large minibatches cause optimization difficulties, but when these are addressed the trained networks exhibit good generalization. Specifically, we show no loss of accuracy when training with large minibatch sizes up to 8192 images. To achieve this result, we adopt a hyper-parameter-free linear scaling rule for adjusting learning rates as a function of minibatch size and develop a new warmup scheme that overcomes optimization challenges early in training. With these simple techniques, our Caffe2-based system trains ResNet-50 with a minibatch size of 8192 on 256 GPUs in one hour, while matching small minibatch accuracy. Using commodity hardware, our implementation achieves ~90% scaling efficiency when moving from 8 to 256 GPUs. Our findings enable training visual recognition models on internet-scale data with high efficiency.