Paper Lineage
Esc
MethodDec 2015arXiv 1512.00567cs.CV

Rethinking the Inception Architecture for Computer Vision

Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe and 2 others

Convolutional networks are at the core of most state-of-the-art computer vision solutions for a wide variety of tasks. Since 2014 very deep convolutional networks started to become mainstream, yielding substantial gains in various benchmarks.

From the abstract

Built on

4 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (4)

  • GoogLeNet (Inception)2014 · cited 7×
    “VGGNet simonyan2014very and GoogLeNet szegedy2015going yielded similarly high performance in the 2014 ILSVRC russakovsky2014imagenet classification challenge.”
    From this paper · §Introduction
  • BatchNorm2015 · cited 4×
    “This is achieved with relatively modest (2.5×2.5\times) increase in computational cost compared to the network described in Ioffe et al ioffe2015batch.”
    From this paper · §Conclusions
  • VGG2014 · cited 2×
    “VGGNet simonyan2014very and GoogLeNet szegedy2015going yielded similarly high performance in the 2014 ILSVRC russakovsky2014imagenet classification challenge.”
    From this paper · §Introduction
  • PReLU / He init2015 · cited 2×
    “Still our solution uses much less computation than the best published results based on denser networks: our model outperforms the results of He et al he2015delving – cutting the top-55 (top-11) error by 25%25\% (14%14\%)…”
    From this paper · §Conclusions

Led to

  • ResNeXt2016 · cited 4×
    “Unlike VGG-nets, the family of Inception models Szegedy2015; Ioffe2015; Szegedy2016a; Szegedy2016 have demonstrated that carefully designed topologies are able to achieve compelling accuracy with low theoretical complexi…”
    From ResNeXt · §Introduction
  • MobileNets2017 · cited 5×, 3 in Method
    “The general trend has been to make deeper and more complicated networks in order to achieve higher accuracy simonyan2014very; szegedy2015rethinking; szegedy2016inception; he2015deep.”
    From MobileNets · §Introduction
  • NASNet2017 · cited 9×, 2 in Method
    “Starting from the seminal work of krizhevsky2012imagenet on using convolutional architectures fukushima1982neocognitron; lecun1998gradient for ImageNet deng2009imagenet classification, successive advancements through arc…”
    From NASNet · §Introduction
  • “Although auxiliary classifier heads were initially proposed to alleviate issues related to vanishing gradients lee2015deeply; szegedy2015going, Szegedy et al. szegedy2016rethinking instead suggest that they also act as r…”
    From Do better ImageNet models transfer bette · §Results
  • EfficientNet2019 · cited 2×, 2 in Method
    “Scaling network depth is the most common way used by many ConvNets (He et al. 2016; Huang et al. 2017; Szegedy et al. 2015; Szegedy et al. 2016).”
    From EfficientNet · §Compound Model Scaling
  • UniCL2022 · cited 2×
    “When fueled with clean and large-scale human-annotated image-label data, e.g., ImageNet deng2009imagenet, supervised learning can attain decent visual recognition capacities over the given categories krizhevsky2012imagen…”
    From UniCL · §Introduction
Abstract

Convolutional networks are at the core of most state-of-the-art computer vision solutions for a wide variety of tasks. Since 2014 very deep convolutional networks started to become mainstream, yielding substantial gains in various benchmarks. Although increased model size and computational cost tend to translate to immediate quality gains for most tasks (as long as enough labeled data is provided for training), computational efficiency and low parameter count are still enabling factors for various use cases such as mobile vision and big-data scenarios. Here we explore ways to scale up networks in ways that aim at utilizing the added computation as efficiently as possible by suitably factorized convolutions and aggressive regularization. We benchmark our methods on the ILSVRC 2012 classification challenge validation set demonstrate substantial gains over the state of the art: 21.2% top-1 and 5.6% top-5 error for single frame evaluation using a network with a computational cost of 5 billion multiply-adds per inference and with using less than 25 million parameters. With an ensemble of 4 models and multi-crop evaluation, we report 3.5% top-5 error on the validation set (3.6% error on the test set) and 17.3% top-1 error on the validation set.