Paper Lineage
Esc
MethodSep 2014arXiv 1409.4842cs.CV

Going Deeper with Convolutions

Christian Szegedy, Wei Liu, Yangqing Jia and 6 others

We propose a deep convolutional neural network architecture codenamed "Inception", which was responsible for setting the new state of the art for classification and detection in the ImageNet Large-Scale Visual Recognition Challenge 2014 (ILSVRC 2014). The main hallmark of this architecture is the improved utilization of the computing resources inside the network.

From the abstract

Built on

2 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (2)

  • OverFeat2013 · cited 3×
    “For larger datasets such as Imagenet, the recent trend has been to increase the number of layers [12] and layer size [21, 14], while using dropout [7] to address the problem of overfitting.”
    From this paper · §Related Work
  • ZFNet2013 · cited 2×
    “Variants of this basic design are prevalent in the image classification literature and have yielded the best results to-date on MNIST, CIFAR and most notably on the ImageNet classification challenge [9, 21].”
    From this paper · §Related Work

Led to

  • VGG2014 · cited 4×, 2 in Method
    “GoogLeNet (Szegedy et al. 2014), a top-performing entry of the ILSVRC-2014 classification task, was developed independently of our work, but is similar in that it is based on very deep ConvNets (22 weight layers) and sma…”
    From VGG · §ConvNet Configurations
  • “We list their performance with a CNN that is equivalent in power (AlexNet krizhevsky2012imagenet) to the one used in this work, though similar to vinyals2014show they outperform our model with a more powerful CNN (VGGNet…”
    From Karpathy visual-semantic alignment · §Experiments
  • PReLU / He init2015 · cited 7×, 1 in Method
    “In Szegedy2014; Lee2014, auxiliary classifiers are added to intermediate layers to help with convergence.”
    From PReLU / He init · §Approach
  • BatchNorm2015 · cited 4×
    “The details of ensemble and multicrop inference are similar to Szegedy et al. 2014.”
    From BatchNorm · §Experiments
  • Ask Your Neurons2015 · cited 2×, 1 in Method
    “In a pilot study, we have found that GoogleNet architecture [11, 29] consistently outperforms the AlexNet architecture [11, 16] as a CNN model for our task and model.”
    From Ask Your Neurons · §Approach
  • FM-IQA2015 · cited 2×, 1 in Method
    “In this paper, we use the GoogleNet [36].”
    From FM-IQA · §The Multimodal QA (mQA) Model
  • “Recent studies have shown that using visual features extracted from convolutional networks trained on large object recognition datasets krizhevsky12; simonyan15; szegedy15 can lead to state-of-the-art results on many vis…”
    From Weakly supervised visual features (Jouli · §Introduction
  • Inception v32015 · cited 7×
    “VGGNet simonyan2014very and GoogLeNet szegedy2015going yielded similarly high performance in the 2014 ILSVRC russakovsky2014imagenet classification challenge.”
    From Inception v3 · §Introduction
  • ResNet2015 · cited 4×
    “In Szegedy2015; Lee2014, a few intermediate layers are directly connected to auxiliary classifiers for addressing vanishing/exploding gradients.”
    From ResNet · §Related Work
  • Wide ResNet2016 · cited 2×
    “Convolutional neural networks have seen a gradual increase of the number of layers in the last few years, starting from AlexNet [Krizhevsky et al.(2012a)Krizhevsky, Sutskever, and Hinton], VGG [Simonyan and Zisserman(201…”
    From Wide ResNet · §Introduction
  • ResNeXt2016 · cited 5×, 1 in Method
    “Unlike VGG-nets, the family of Inception models Szegedy2015; Ioffe2015; Szegedy2016a; Szegedy2016 have demonstrated that carefully designed topologies are able to achieve compelling accuracy with low theoretical complexi…”
    From ResNeXt · §Introduction
  • Goyal large-batch SGD2017 · cited 3×
    “We use scale and aspect ratio data augmentation Szegedy2015 as in Gross2016.”
    From Goyal large-batch SGD · §Main Results and Analysis
  • NASNet2017 · cited 5×, 2 in Method
    “Starting from the seminal work of krizhevsky2012imagenet on using convolutional architectures fukushima1982neocognitron; lecun1998gradient for ImageNet deng2009imagenet classification, successive advancements through arc…”
    From NASNet · §Introduction
  • “Although auxiliary classifier heads were initially proposed to alleviate issues related to vanishing gradients lee2015deeply; szegedy2015going, Szegedy et al. szegedy2016rethinking instead suggest that they also act as r…”
    From Do better ImageNet models transfer bette · §Results
  • ImageNet-trained CNNs are biased towards2018 · cited 1×, 1 in Method
    “The same images were fed to four CNNs pre-trained on standard ImageNet, namely AlexNet (Krizhevsky et al. 2012), GoogLeNet (Szegedy et al. 2015), VGG-16 (Simonyan & Zisserman 2015) and ResNet-50 (He et al. 2015).”
    From ImageNet-trained CNNs are biased towards · §Methods
  • EfficientNet2019 · cited 2×, 1 in Method
    “Scaling network depth is the most common way used by many ConvNets (He et al. 2016; Huang et al. 2017; Szegedy et al. 2015; Szegedy et al. 2016).”
    From EfficientNet · §Compound Model Scaling
  • Noisy Student2019 · cited 2×
    “Deep learning has shown remarkable successes in image recognition in recent years krizhevsky2012imagenet; szegedy2015going; simonyan2014very; he2016deep; tan2019efficientnet.”
    From Noisy Student · §Introduction
  • SimCLR2020 · cited 2×
    “The other type of augmentation involves appearance transformation, such as color distortion (including color dropping, brightness, contrast, saturation, hue) (Howard 2013; Szegedy et al. 2015), Gaussian blur, and Sobel f…”
    From SimCLR · §Data Augmentation for Contrastive Representation Learning
  • Natural distribution shift robustness2020 · cited 1×, 1 in Method
    “This category includes 78 models with architectures ranging from AlexNet to EfficietNet, e.g., [50, 78, 37, 85, 88].”
    From Natural distribution shift robustness · §Experimental setup
  • Perceiver2021 · cited 2×
    “ImageNet has been a crucial bellwether in the development of architectures for image recognition (Krizhevsky et al. 2012; Simonyan & Zisserman 2015; Szegedy et al. 2015; He et al. 2016) and, until recently, it has been d…”
    From Perceiver · §Experiments
Abstract

We propose a deep convolutional neural network architecture codenamed "Inception", which was responsible for setting the new state of the art for classification and detection in the ImageNet Large-Scale Visual Recognition Challenge 2014 (ILSVRC 2014). The main hallmark of this architecture is the improved utilization of the computing resources inside the network. This was achieved by a carefully crafted design that allows for increasing the depth and width of the network while keeping the computational budget constant. To optimize quality, the architectural decisions were based on the Hebbian principle and the intuition of multi-scale processing. One particular incarnation used in our submission for ILSVRC 2014 is called GoogLeNet, a 22 layers deep network, the quality of which is assessed in the context of classification and detection.