Aggregated Residual Transformations for Deep Neural Networks
We present a simple, highly modularized network architecture for image classification. Our network is constructed by repeating a building block that aggregates a set of transformations with the same topology.
Also cited · not yet reviewed (8)
- ResNet2015 · cited 18×, 7 in Method“ResNets He2016 can be thought of as two-branch networks where one branch is the identity mapping.”From this paper · §Related Work
- GoogLeNet (Inception)2014 · cited 5×, 1 in Method“Unlike VGG-nets, the family of Inception models Szegedy2015; Ioffe2015; Szegedy2016a; Szegedy2016 have demonstrated that carefully designed topologies are able to achieve compelling accuracy with low theoretical complexi…”From this paper · §Introduction
- BatchNorm2015 · cited 3×, 1 in Method“Unlike VGG-nets, the family of Inception models Szegedy2015; Ioffe2015; Szegedy2016a; Szegedy2016 have demonstrated that carefully designed topologies are able to achieve compelling accuracy with low theoretical complexi…”From this paper · §Introduction
- PReLU / He init2015 · cited 1×, 1 in Method“We adopt the weight initialization of He2015.”From this paper · §Implementation details
Show 4 more
- Inception v32015 · cited 4דUnlike VGG-nets, the family of Inception models Szegedy2015; Ioffe2015; Szegedy2016a; Szegedy2016 have demonstrated that carefully designed topologies are able to achieve compelling accuracy with low theoretical complexi…”From this paper · §Introduction
- MS COCO2014 · cited 3דThis paper further evaluates ResNeXt on a larger ImageNet-5K set and the COCO object detection dataset Lin2014, showing consistently better accuracy than its ResNet counterparts.”From this paper · §Introduction
- DeCAF2013 · cited 2דIn contrast to traditional hand-designed features (e.g., SIFT Lowe2004 and HOG Dalal2005), features learned by neural networks from large-scale data Russakovsky2015 require minimal human involvement during training, and…”From this paper · §Introduction
- VGG2014 · cited 2דThe VGG-nets Simonyan2015 exhibit a simple yet effective strategy of constructing very deep networks: stacking building blocks of the same shape.”From this paper · §Introduction
Led to
- Goyal large-batch SGD2017 · cited 2דTo investigate this question, we adopt the ImageNet-5k dataset suggested by Xie et al. xie2017 that extends ImageNet-1k to 6.8 million images (roughly 5×\times larger) by adding 4k additional categories from ImageNet-22k…”From Goyal large-batch SGD · §Main Results and Analysis
- NASNet2017 · cited 2דStarting from the seminal work of krizhevsky2012imagenet on using convolutional architectures fukushima1982neocognitron; lecun1998gradient for ImageNet deng2009imagenet classification, successive advancements through arc…”From NASNet · §Introduction
- MobileNetV22018 · cited 2×, 1 in Method“When the expansion ratio is smaller than 11, this is a classical residual convolutional block ResNet; ResNext2016.”From MobileNetV2 · §Preliminaries, discussion and intuition
- Instagram hashtag pre-training2018 · cited 3×, 2 in Method“For the 5k set, we use the now standard IN-5k proposed in [15] (6.6M training images).”From Instagram hashtag pre-training · §Scaling up Supervised Pretraining
- ImageNet-C2019 · cited 2דUnlike current deep learning classifiers (Krizhevsky et al. 2012; He et al. 2015; Xie et al. 2016), the human vision system is not fooled by small changes in query images.”From ImageNet-C · §Introduction
- Billion-scale semi-supervised2019 · cited 3דModels: For student and teacher models, we use residual networks [16], ResNet-d with dd = {18,50}\{18,50\} and residual networks with group convolutions [43], ResNeXt-101 32XCd with 101101 layers and group widths CC = {4…”From Billion-scale semi-supervised · §Image classification: experiments & analysis
- VL-BERT meta-analysis2020 · cited 1×, 1 in Method“V&L BERTs typically extract image features using a Faster R-CNN (Ren et al. 2015) trained on the Visual Genome dataset (VG; Krishna et al. 2017), either with a ResNet-101101 (He et al. 2016) or a ResNeXT-152152 backbone…”From VL-BERT meta-analysis · §Experimental Setup
- Swin2021 · cited 3דIn addition to these architectural advances, there has also been much work on improving individual convolution layers, such as depth-wise convolution [70] and deformable convolution [18, 84].”From Swin · §Related Work
Abstract
We present a simple, highly modularized network architecture for image classification. Our network is constructed by repeating a building block that aggregates a set of transformations with the same topology. Our simple design results in a homogeneous, multi-branch architecture that has only a few hyper-parameters to set. This strategy exposes a new dimension, which we call "cardinality" (the size of the set of transformations), as an essential factor in addition to the dimensions of depth and width. On the ImageNet-1K dataset, we empirically show that even under the restricted condition of maintaining complexity, increasing cardinality is able to improve classification accuracy. Moreover, increasing cardinality is more effective than going deeper or wider when we increase the capacity. Our models, named ResNeXt, are the foundations of our entry to the ILSVRC 2016 classification task in which we secured 2nd place. We further investigate ResNeXt on an ImageNet-5K set and the COCO detection set, also showing better results than its ResNet counterpart. The code and models are publicly available online.