Paper Lineage
Esc
MethodApr 2017arXiv 1704.04861cs.CV

MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications

Andrew G. Howard, Menglong Zhu, Bo Chen and 5 others

We present a class of efficient models called MobileNets for mobile and embedded vision applications. MobileNets are based on a streamlined architecture that uses depth-wise separable convolutions to build light weight deep neural networks.

From the abstract

Built on

4 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (4)

  • Inception v32015 · cited 5×, 3 in Method
    “The general trend has been to make deeper and more complicated networks in order to achieve higher accuracy simonyan2014very; szegedy2015rethinking; szegedy2016inception; he2015deep.”
    From this paper · §Introduction
  • BatchNorm2015 · cited 3×, 1 in Method
    “MobileNets are built primarily from depthwise separable convolutions initially introduced in sifre2014rigid and subsequently used in Inception models ioffe2015batch to reduce the computation in the first few layers.”
    From this paper · §Prior Work
  • Knowledge Distillation2015 · cited 3×
    “Another method for training small networks is distillation hinton2015distilling which uses a larger network to teach a smaller network.”
    From this paper · §Prior Work
  • VGG2014 · cited 2×
    “The general trend has been to make deeper and more complicated networks in order to achieve higher accuracy simonyan2014very; szegedy2015rethinking; szegedy2016inception; he2015deep.”
    From this paper · §Introduction

Led to

  • NASNet2017 · cited 5×
    “Notably, the smallest version of NASNet achieves 74.0% top-1 accuracy on ImageNet, which is 3.1% better than previously engineered architectures targeted towards mobile and embedded vision tasks howard2017mobilenets; shu…”
    From NASNet · §Introduction
  • MobileNetV22018 · cited 8×, 6 in Method
    “Our network design is based on MobileNetV1 MobilenetV1.”
    From MobileNetV2 · §Related Work
  • “Chatfield14; simonyan2014very; huang2016; howard2017mobilenets; he2017mask), they have never been systematically explored across network architectures.”
    From Do better ImageNet models transfer bette · §Introduction
  • MnasNet2018 · cited 5×, 1 in Method
    “Given restricted computational resources available on mobile devices, much recent research has focused on designing and improving mobile CNN models by reducing the depth of the network and utilizing less expensive operat…”
    From MnasNet · §Introduction
  • EfficientNet2019 · cited 5×, 1 in Method
    “Scaling network width is commonly used for small size models (Howard et al. 2017; Sandler et al. 2018; Tan et al. 2019)22 2 In some literature, scaling number of channels is called “depth multiplier”, which means the sam…”
    From EfficientNet · §Compound Model Scaling
Abstract

We present a class of efficient models called MobileNets for mobile and embedded vision applications. MobileNets are based on a streamlined architecture that uses depth-wise separable convolutions to build light weight deep neural networks. We introduce two simple global hyper-parameters that efficiently trade off between latency and accuracy. These hyper-parameters allow the model builder to choose the right sized model for their application based on the constraints of the problem. We present extensive experiments on resource and accuracy tradeoffs and show strong performance compared to other popular models on ImageNet classification. We then demonstrate the effectiveness of MobileNets across a wide range of applications and use cases including object detection, finegrain classification, face attributes and large scale geo-localization.