Very Deep Convolutional Networks for Large-Scale Image Recognition
In this work we investigate the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting. Our main contribution is a thorough evaluation of networks of increasing depth using an architecture with very small (3x3) convolution filters, which shows that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 weight layers.
Also cited · not yet reviewed (3)
- OverFeat2013 · cited 9×, 2 in Method“Convolutional networks (ConvNets) have recently enjoyed a great success in large-scale image and video recognition (Krizhevsky et al. 2012; Zeiler & Fergus 2013; Sermanet et al. 2014; Simonyan & Zisserman 2014) which has…”From this paper · §Introduction
- GoogLeNet (Inception)2014 · cited 4×, 2 in Method“GoogLeNet (Szegedy et al. 2014), a top-performing entry of the ILSVRC-2014 classification task, was developed independently of our work, but is similar in that it is based on very deep ConvNets (22 weight layers) and sma…”From this paper · §ConvNet Configurations
- ZFNet2013 · cited 6×, 1 in Method“In our experiments, we evaluated models trained at two fixed scales: S=256S=256 (which has been widely used in the prior art (Krizhevsky et al. 2012; Zeiler & Fergus 2013; Sermanet et al. 2014)) and S=384S=384.”From this paper · §Classification Framework
Led to
- Captions to Visual Concepts2014 · cited 5דNext, we featurize each of these regions using rich convolutional neural network (CNN) features, fine-tuned on our training data krizhevskyNIPS12; simonyan14very.”From Captions to Visual Concepts · §Introduction
- Karpathy visual-semantic alignment2014 · cited 3דWe list their performance with a CNN that is equivalent in power (AlexNet krizhevsky2012imagenet) to the one used in this work, though similar to vinyals2014show they outperform our model with a more powerful CNN (VGGNet…”From Karpathy visual-semantic alignment · §Experiments
- PReLU / He init2015 · cited 21×, 14 in Method“With fixed standard deviations (e.g., 0.01 in Krizhevsky2012), very deep models (e.g., >>8 conv layers) have difficulties to converge, as reported by the VGG team Simonyan2014 and also observed in our experiments.”From PReLU / He init · §Approach
- Learning like a Child2015 · cited 2×, 1 in Method“The vision component contains a 16-layer deep convolutional neural network (CNN [48]) pre-trained on the ImageNet classification task [45].”From Learning like a Child · §The Image Captioning Model
- VQA2015 · cited 3×, 3 in Method“I: The activations from the last hidden layer of VGGNet [48] are used as 4096-dim image embedding.”From VQA · §V VQA Baselines and Methods
- Flickr30k Entities2015 · cited 2דBy using better region features (Fast RCNN (Girshick, 2015) instead of ImageNet-trained VGG (Simonyan and Zisserman, 2014)) in combination with size and color cues, we are able to improve the Recall@1 for phrase localiza…”From Flickr30k Entities · §Introduction
- FM-IQA2015 · cited 2×, 1 in Method“For computer vision, methods based on Convolutional Neural Network (CNN [20]) achieve the state-of-the-art performance in various tasks, such as object classification [17, 34, 17], detection [10, 44] and segmentation [3]…”From FM-IQA · §Related Work
- Inception v32015 · cited 2דVGGNet simonyan2014very and GoogLeNet szegedy2015going yielded similarly high performance in the 2014 ILSVRC russakovsky2014imagenet classification challenge.”From Inception v3 · §Introduction
- ResNet2015 · cited 11דOur plain baselines (Fig. 3, middle) are mainly inspired by the philosophy of VGG nets Simonyan2015 (Fig. 3, left).”From ResNet · §Deep Residual Learning
- Visual Genome2016 · cited 5דMuch progress has been made in recent years towards this goal, including image classification Deng et al., 2009; Perronnin et al., 2010; Simonyan and Zisserman, 2014; Krizhevsky et al., 2012; Szegedy et al., 2014 and obj…”From Visual Genome · §Introduction
- Wide ResNet2016 · cited 4דConvolutional neural networks have seen a gradual increase of the number of layers in the last few years, starting from AlexNet [Krizhevsky et al.(2012a)Krizhevsky, Sutskever, and Hinton], VGG [Simonyan and Zisserman(201…”From Wide ResNet · §Introduction
- ResNeXt2016 · cited 2דThe VGG-nets Simonyan2015 exhibit a simple yet effective strategy of constructing very deep networks: stacking building blocks of the same shape.”From ResNeXt · §Introduction
- Constrained beam search captioning2016 · cited 1×, 1 in Method“For the CNN component of the model, we evaluate using the 16-layer VGG Simonyan and Zisserman 2015 model and the 50-layer Residual Net He et al. 2016, pretrained on ILSVRC-2012 Russakovsky et al. 2015 in both cases.”From Constrained beam search captioning · §Approach
- VQA v22016 · cited 3×, 1 in Method“It should be noted that MCB uses image features from a more powerful CNN architecture ResNet ResNet while the previous two models use image features from VGGNet Simonyan15.”From VQA v2 · §Benchmarking Existing VQA Models
- MobileNets2017 · cited 2דThe general trend has been to make deeper and more complicated networks in order to achieve higher accuracy simonyan2014very; szegedy2015rethinking; szegedy2016inception; he2015deep.”From MobileNets · §Introduction
- Goyal large-batch SGD2017 · cited 2דWe are in an unprecedented era in AI research history in which the increasing data and model scale is rapidly improving accuracy in computer vision Krizhevsky2012; Zeiler2014; Sermanet2014; Simonyan2015; Szegedy2015; He2…”From Goyal large-batch SGD · §Introduction
- NASNet2017 · cited 4×, 1 in Method“Starting from the seminal work of krizhevsky2012imagenet on using convolutional architectures fukushima1982neocognitron; lecun1998gradient for ImageNet deng2009imagenet classification, successive advancements through arc…”From NASNet · §Introduction
- MobileNetV22018 · cited 2דBoth manual architecture search and improvements in training algorithms, carried out by numerous teams has lead to dramatic improvements over early designs such as AlexNet AlexNet, VGGNet VGGNet, GoogLeNet GoogleNet. , a…”From MobileNetV2 · §Related Work
- Instance discrimination2018 · cited 2דAs the network architecture has a big impact on the performance, we consider a few typical architectures: AlexNet krizhevsky2012imagenet, VGG16 simonyan2014very, ResNet-18, and ResNet-50 he2015deep.”From Instance discrimination · §Experiments
- ImageNet-trained CNNs are biased towards2018 · cited 1×, 1 in Method“The same images were fed to four CNNs pre-trained on standard ImageNet, namely AlexNet (Krizhevsky et al. 2012), GoogLeNet (Szegedy et al. 2015), VGG-16 (Simonyan & Zisserman 2015) and ResNet-50 (He et al. 2015).”From ImageNet-trained CNNs are biased towards · §Methods
- ImageNetV22019 · cited 2דThe models include the seminal AlexNet [36], widely used convolutional networks [49, 21, 27, 52], and the state-of-the-art [8, 39].”From ImageNetV2 · §Summary of Our Experiments
- Grid features for VQA2020 · cited 2דAs a result, images are represented by a collection of bounding box or region11 1 We use the terms ‘region’ and ‘bounding box’ interchangeably.-based features anderson2018bottom; teney2018tips–in contrast to vanilla grid…”From Grid features for VQA · §Introduction
- BYOL2020 · cited 2ד[1, 2, 3, 4, 5, 6, 7]”From BYOL · §Introduction
- Natural distribution shift robustness2020 · cited 1×, 1 in Method“This category includes 78 models with architectures ranging from AlexNet to EfficietNet, e.g., [50, 78, 37, 85, 88].”From Natural distribution shift robustness · §Experimental setup
- Swin2021 · cited 2×, 1 in Method“These stages jointly produce a hierarchical representation, with the same feature map resolutions as those of typical convolutional networks, e.g., VGG [52] and ResNet [30].”From Swin · §Method
- Swin V22021 · cited 2דWhile it has long been recognized that larger vision models usually perform better on vision tasks simonyan2014vgg; he2015resnet, the absolute model size was just able to reach about 1-2 billion parameters very recently…”From Swin V2 · §Introduction
Abstract
In this work we investigate the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting. Our main contribution is a thorough evaluation of networks of increasing depth using an architecture with very small (3x3) convolution filters, which shows that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 weight layers. These findings were the basis of our ImageNet Challenge 2014 submission, where our team secured the first and the second places in the localisation and classification tracks respectively. We also show that our representations generalise well to other datasets, where they achieve state-of-the-art results. We have made our two best-performing ConvNet models publicly available to facilitate further research on the use of deep visual representations in computer vision.