Paper Lineage
Esc
MethodMay 2019arXiv 1905.11946cs.LG

EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks

Mingxing Tan, Quoc V. Le

Convolutional Neural Networks (ConvNets) are commonly developed at a fixed resource budget, and then scaled up for better accuracy if more resources are available. In this paper, we systematically study model scaling and identify that carefully balancing network depth, width, and resolution can lead to better performance.

From the abstract

Built on

10 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (10)

  • MnasNet2018 · cited 10×, 6 in Method
    “Notably, the effectiveness of model scaling heavily depends on the baseline network; to go even further, we use neural architecture search (Zoph & Le 2017; Tan et al. 2019) to develop a new baseline network, and scale it…”
    From this paper · §Introduction
  • ResNet2015 · cited 11×, 3 in Method
    “While these models are mainly designed for ImageNet, recent studies have shown better ImageNet models also perform better across a variety of transfer learning datasets (Kornblith et al. 2019), and other computer vision…”
    From this paper · §Related Work
  • Wide ResNet2016 · cited 5×, 2 in Method
    “However, deeper networks are also more difficult to train due to the vanishing gradient problem (Zagoruyko & Komodakis 2016).”
    From this paper · §Compound Model Scaling
  • MobileNetV22018 · cited 5×, 2 in Method
    “Scaling network width is commonly used for small size models (Howard et al. 2017; Sandler et al. 2018; Tan et al. 2019)22 2 In some literature, scaling number of channels is called “depth multiplier”, which means the sam…”
    From this paper · §Compound Model Scaling
Show 6 more
  • NASNet2017 · cited 4×, 2 in Method
    “Table 5 shows the transfer learning performance: (1) Compared to public available models, such as NASNet-A (Zoph et al. 2018) and Inception-v4 (Szegedy et al. 2017), our EfficientNet models achieve better accuracy with 4…”
    From this paper · §Experiments
  • Inception v32015 · cited 2×, 2 in Method
    “Scaling network depth is the most common way used by many ConvNets (He et al. 2016; Huang et al. 2017; Szegedy et al. 2015; Szegedy et al. 2016).”
    From this paper · §Compound Model Scaling
  • MobileNets2017 · cited 5×, 1 in Method
    “Scaling network width is commonly used for small size models (Howard et al. 2017; Sandler et al. 2018; Tan et al. 2019)22 2 In some literature, scaling number of channels is called “depth multiplier”, which means the sam…”
    From this paper · §Compound Model Scaling
  • GoogLeNet (Inception)2014 · cited 2×, 1 in Method
    “Scaling network depth is the most common way used by many ConvNets (He et al. 2016; Huang et al. 2017; Szegedy et al. 2015; Szegedy et al. 2016).”
    From this paper · §Compound Model Scaling
  • BatchNorm2015 · cited 1×, 1 in Method
    “Although several techniques, such as skip connections (He et al. 2016) and batch normalization (Ioffe & Szegedy 2015), alleviate the training problem, the accuracy gain of very deep network diminishes: for example, ResNe…”
    From this paper · §Compound Model Scaling
  • “While these models are mainly designed for ImageNet, recent studies have shown better ImageNet models also perform better across a variety of transfer learning datasets (Kornblith et al. 2019), and other computer vision…”
    From this paper · §Related Work

Led to

  • Noisy Student2019 · cited 6×
    “Deep learning has shown remarkable successes in image recognition in recent years krizhevsky2012imagenet; szegedy2015going; simonyan2014very; he2016deep; tan2019efficientnet.”
    From Noisy Student · §Introduction
  • Kaplan scaling laws2020 · cited 2×
    “EfficientNets [TL19] also appear to obey an approximate power-law relation between accuracy and model size.”
    From Kaplan scaling laws · §Related Work
  • Natural distribution shift robustness2020 · cited 1×, 1 in Method
    “This category includes 78 models with architectures ranging from AlexNet to EfficietNet, e.g., [50, 78, 37, 85, 88].”
    From Natural distribution shift robustness · §Experimental setup
  • DeiT2020 · cited 2×
    “The evolution of the state of the art on the ImageNet dataset [42] reflects the progress with convolutional neural network architectures and learning [32, 44, 48, 50, 51, 57].”
    From DeiT · §Related work
  • CLIP2021 · cited 2×, 2 in Method
    “While previous computer vision research has often scaled models by increasing the width (Mahajan et al. 2018) or depth (He et al. 2016a) in isolation, for the ResNet image encoders we adapt the approach of Tan & Le 2019…”
    From CLIP · §Approach
  • Swin2021 · cited 3×
    “Since then, deeper and more effective convolutional neural architectures have been proposed to further propel the deep learning wave in computer vision, e.g., VGG [52], GoogleNet [57], ResNet [30], DenseNet [34], HRNet […”
    From Swin · §Related Work
  • CoAtNet2021 · cited 3×
    “Recent works show that an improved inverted residual bottlenecks (MBConv [27, 35]), which is built upon depthwise convolutions, can achieve both high accuracy and better efficiency [5, 19].”
    From CoAtNet · §Related Work
Abstract

Convolutional Neural Networks (ConvNets) are commonly developed at a fixed resource budget, and then scaled up for better accuracy if more resources are available. In this paper, we systematically study model scaling and identify that carefully balancing network depth, width, and resolution can lead to better performance. Based on this observation, we propose a new scaling method that uniformly scales all dimensions of depth/width/resolution using a simple yet highly effective compound coefficient. We demonstrate the effectiveness of this method on scaling up MobileNets and ResNet. To go even further, we use neural architecture search to design a new baseline network and scale it up to obtain a family of models, called EfficientNets, which achieve much better accuracy and efficiency than previous ConvNets. In particular, our EfficientNet-B7 achieves state-of-the-art 84.3% top-1 accuracy on ImageNet, while being 8.4x smaller and 6.1x faster on inference than the best existing ConvNet. Our EfficientNets also transfer well and achieve state-of-the-art accuracy on CIFAR-100 (91.7%), Flowers (98.8%), and 3 other transfer learning datasets, with an order of magnitude fewer parameters. Source code is at https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet.