Paper Lineage
Esc
MethodDec 2015arXiv 1512.03385cs.CV

Deep Residual Learning for Image Recognition

Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun

Deeper neural networks are more difficult to train. We present a residual learning framework to ease the training of networks that are substantially deeper than those used previously.

From the abstract

Built on

5 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (5)

  • VGG2014 · cited 11×
    “Our plain baselines (Fig. 3, middle) are mainly inspired by the philosophy of VGG nets Simonyan2015 (Fig. 3, left).”
    From this paper · §Deep Residual Learning
  • BatchNorm2015 · cited 8×
    “This problem, however, has been largely addressed by normalized initialization LeCun1998; Glorot2010; Saxe2013; He2015 and intermediate normalization layers Ioffe2015, which enable networks with tens of layers to start c…”
    From this paper · §Introduction
  • PReLU / He init2015 · cited 5×
    “This problem, however, has been largely addressed by normalized initialization LeCun1998; Glorot2010; Saxe2013; He2015 and intermediate normalization layers Ioffe2015, which enable networks with tens of layers to start c…”
    From this paper · §Introduction
  • GoogLeNet (Inception)2014 · cited 4×
    “In Szegedy2015; Lee2014, a few intermediate layers are directly connected to auxiliary classifiers for addressing vanishing/exploding gradients.”
    From this paper · §Related Work
Show 1 more
  • Faster R-CNN2015 · cited 2×
    “We adopt Faster R-CNN Ren2015 as the detection method.”
    From this paper · §Experiments

Led to

  • Faster R-CNN2015 · cited 7×
    “In ILSVRC and COCO 2015 competitions, Faster R-CNN and RPN are the basis of several 1st-place entries He2015a in the tracks of ImageNet detection, ImageNet localization, COCO detection, and COCO segmentation.”
    From Faster R-CNN · §I Introduction
  • Wide ResNet2016 · cited 3×
    “Convolutional neural networks have seen a gradual increase of the number of layers in the last few years, starting from AlexNet [Krizhevsky et al.(2012a)Krizhevsky, Sutskever, and Hinton], VGG [Simonyan and Zisserman(201…”
    From Wide ResNet · §Introduction
  • GNMT2016 · cited 2×, 1 in Method
    “Motivated by the idea of modeling differences between an intermediate layer’s output and the targets, which has shown to work well for many projects in the past [16, 21, 40], we introduce residual connections among the L…”
    From GNMT · §Model Architecture
  • ResNeXt2016 · cited 18×, 7 in Method
    “ResNets He2016 can be thought of as two-branch networks where one branch is the identity mapping.”
    From ResNeXt · §Related Work
  • Constrained beam search captioning2016 · cited 3×, 1 in Method
    “For this task we use the ResNet-50 He et al. 2016 CNN, and train the base model on a combined training set containing 155k images comprised of the MSCOCO Chen et al. 2015 training and validation datasets, and the full Fl…”
    From Constrained beam search captioning · §Experiments
  • VQA v22016 · cited 1×, 1 in Method
    “It should be noted that MCB uses image features from a more powerful CNN architecture ResNet ResNet while the previous two models use image features from VGGNet Simonyan15.”
    From VQA v2 · §Benchmarking Existing VQA Models
  • Goyal large-batch SGD2017 · cited 20×
    “Of equal importance, in a research domain, we have found it to simplify migrating algorithms from a single-GPU to a multi-GPU implementation without requiring hyper-parameter search, e.g. in our experience migrating Fast…”
    From Goyal large-batch SGD · §Introduction
  • Transformer2017 · cited 1×, 1 in Method
    “We employ a residual connection [11] around each of the two sub-layers, followed by layer normalization [1].”
    From Transformer · §Model Architecture
  • JFT-300M (unreasonable effectiveness)2017 · cited 6×, 3 in Method
    “Our baseline ResNet-101 performs 1% better than the open-sourced ResNet-101 checkpoint from the authors of [16], using the same evaluation protocol.”
    From JFT-300M (unreasonable effectiveness) · §Training and Evaluation Framework
  • NASNet2017 · cited 5×, 2 in Method
    “Starting from the seminal work of krizhevsky2012imagenet on using convolutional architectures fukushima1982neocognitron; lecun1998gradient for ImageNet deng2009imagenet classification, successive advancements through arc…”
    From NASNet · §Introduction
  • Bottom-Up Top-Down attention2017 · cited 2×, 1 in Method
    “In this work, we use Faster R-CNN in conjunction with the ResNet-101 he2015deep CNN.”
    From Bottom-Up Top-Down attention · §Approach
  • MobileNetV22018 · cited 5×, 2 in Method
    “Both manual architecture search and improvements in training algorithms, carried out by numerous teams has lead to dramatic improvements over early designs such as AlexNet AlexNet, VGGNet VGGNet, GoogLeNet GoogleNet. , a…”
    From MobileNetV2 · §Related Work
  • Neural Baby Talk2018 · cited 2×, 2 in Method
    “We use Faster R-CNN ren2015faster with ResNet-101 he2015deep to obtain region proposals for the image.”
    From Neural Baby Talk · §Method
  • Instagram hashtag pre-training2018 · cited 1×, 1 in Method
    “We believe our results will generalize to other architectures [24, 25, 26].”
    From Instagram hashtag pre-training · §Scaling up Supervised Pretraining
  • Instance discrimination2018 · cited 2×
    “As the network architecture has a big impact on the performance, we consider a few typical architectures: AlexNet krizhevsky2012imagenet, VGG16 simonyan2014very, ResNet-18, and ResNet-50 he2015deep.”
    From Instance discrimination · §Experiments
  • Box Attention2018 · cited 3×
    “In our experiments we use the RetinaNet [19] model with ResNet50 [11] backbone as the base detection model.”
    From Box Attention · §Experiments
  • MnasNet2018 · cited 2×
    “Compared to the widely used ResNet-50 resnet16, our MnasNet model achieves slightly higher (76.7%) accuracy with 4.8×4.8\times fewer parameters and 𝟏𝟎×10\times fewer multiply-add operations.”
    From MnasNet · §Introduction
  • ImageNet-trained CNNs are biased towards2018 · cited 1×, 1 in Method
    “The same images were fed to four CNNs pre-trained on standard ImageNet, namely AlexNet (Krizhevsky et al. 2012), GoogLeNet (Szegedy et al. 2015), VGG-16 (Simonyan & Zisserman 2015) and ResNet-50 (He et al. 2015).”
    From ImageNet-trained CNNs are biased towards · §Methods
  • ImageNetV22019 · cited 2×
    “The models include the seminal AlexNet [36], widely used convolutional networks [49, 21, 27, 52], and the state-of-the-art [8, 39].”
    From ImageNetV2 · §Summary of Our Experiments
  • “Models: For student and teacher models, we use residual networks [16], ResNet-d with dd = {18,50}\{18,50\} and residual networks with group convolutions [43], ResNeXt-101 32XCd with 101101 layers and group widths CC = {4…”
    From Billion-scale semi-supervised · §Image classification: experiments & analysis
  • CPC v22019 · cited 1×, 1 in Method
    “For this task, the classifier hψh_{\psi} is an arbitrary deep neural network (we use an 11-block ResNet architecture (He et al. 2016a) with 4096-dimensional feature maps and 1024-dimensional bottleneck layers).”
    From CPC v2 · §Experimental Setup
  • EfficientNet2019 · cited 11×, 3 in Method
    “While these models are mainly designed for ImageNet, recent studies have shown better ImageNet models also perform better across a variety of transfer learning datasets (Kornblith et al. 2019), and other computer vision…”
    From EfficientNet · §Related Work
  • AMDIM2019 · cited 1×, 1 in Method
    “Our model uses an encoder based on the standard ResNet (He et al. 2016a; He et al. 2016b), with changes to make it suitable for DIM.”
    From AMDIM · §Method Description
  • ViLBERT2019 · cited 2×
    “We use Faster R-CNN [31] (with ResNet-101 [11] backbone) pretrained on the Visual Genome dataset [16] (see [30] for details) to extract region features.”
    From ViLBERT · §Experimental Settings
  • Decoupled box proposals captioning2019 · cited 1×, 1 in Method
    “ResNet-101 He et al. 2016 pre-trained on ImageNet Russakovsky et al. 2015 is used as the core featurization network11 1 See further details in the supplementary material..”
    From Decoupled box proposals captioning · §Features and Experimental Setup
  • T52019 · cited 1×, 1 in Method
    “After layer normalization, a residual skip connection (He et al. 2016) adds each subcomponent’s input to its output.”
    From T5 · §Setup
  • Noisy Student2019 · cited 2×
    “Deep learning has shown remarkable successes in image recognition in recent years krizhevsky2012imagenet; szegedy2015going; simonyan2014very; he2016deep; tan2019efficientnet.”
    From Noisy Student · §Introduction
  • MoCo2019 · cited 4×, 2 in Method
    “We adopt a ResNet He2016 as the encoder, whose last fully-connected layer (after global average pooling) has a fixed-dimensional output (128-D Wu2018a).”
    From MoCo · §Method
  • Grid features for VQA2020 · cited 6×
    “As a result, images are represented by a collection of bounding box or region11 1 We use the terms ‘region’ and ‘bounding box’ interchangeably.-based features anderson2018bottom; teney2018tips–in contrast to vanilla grid…”
    From Grid features for VQA · §Introduction
  • SimCLR2020 · cited 2×, 2 in Method
    “We opt for simplicity and adopt the commonly used ResNet (He et al. 2016) to obtain 𝒉i=f⁡(𝒙~i)=ResNet⁡(𝒙~i)\bm{h}_{i}=f(\tilde{\bm{x}}_{i})=ResNet(\tilde{\bm{x}}_{i}) where 𝒉i∈ℝd\bm{h}_{i}\in\mathbb{R}^{d} is the out…”
    From SimCLR · §Method
  • BYOL2020 · cited 2×, 1 in Method
    “We use a convolutional residual network [22] with 50 layers and post-activation (ResNet-50(1×)50(1\times) v1) as our base parametric encoders fθf_{\theta} and fξf_{\xi}.”
    From BYOL · §Method
  • GShard2020 · cited 2×
    “For years, the fields have been continuously reporting new state of the art results using varieties of model architectures for computer vision tasks [57, 58, 7], for natural language understanding tasks [59, 60, 61], for…”
    From GShard · §Related Work
  • Natural distribution shift robustness2020 · cited 1×, 1 in Method
    “This category includes 78 models with architectures ranging from AlexNet to EfficietNet, e.g., [50, 78, 37, 85, 88].”
    From Natural distribution shift robustness · §Experimental setup
  • ConVIRT2020 · cited 1×, 1 in Method
    “For the image encoder fvf_{v}, we use the ResNet50 architecture He et al. 2016 for all experiments, as it is the architecture of choice for much medical imaging work and is shown to achieve competitive performance.”
    From ConVIRT · §Methods
  • ViT2020 · cited 2×
    “In computer vision, however, convolutional architectures remain dominant (LeCun et al. 1989; Krizhevsky et al. 2012; He et al. 2016).”
    From ViT · §Introduction
  • VL-BERT meta-analysis2020 · cited 1×, 1 in Method
    “V&L BERTs typically extract image features using a Faster R-CNN (Ren et al. 2015) trained on the Visual Genome dataset (VG; Krishna et al. 2017), either with a ResNet-101101 (He et al. 2016) or a ResNeXT-152152 backbone…”
    From VL-BERT meta-analysis · §Experimental Setup
  • Conceptual 12M2021 · cited 1×, 1 in Method
    “We train a Faster-RCNN [68] on Visual Genome [43], with a ResNet101 [29] backbone trained on JFT [30] and fine-tuned on ImageNet [69].”
    From Conceptual 12M · §Evaluating Vision-and-Language Pre-Training Data
  • CLIP2021 · cited 2×, 2 in Method
    “For the first, we use ResNet-50 (He et al. 2016a) as the base architecture for the image encoder due to its widespread adoption and proven performance.”
    From CLIP · §Approach
  • Perceiver2021 · cited 3×, 1 in Method
    “In contrast, the ConvNets that are typically used in image processing – such as residual networks (ResNets) (He et al. 2016) – bake in 2D spatial structure in several ways, including by using filters that look only at lo…”
    From Perceiver · §Methods
  • Swin2021 · cited 4×, 1 in Method
    “These stages jointly produce a hierarchical representation, with the same feature map resolutions as those of typical convolutional networks, e.g., VGG [52] and ResNet [30].”
    From Swin · §Method
  • CoAtNet2021 · cited 2×
    “Traditionally, regular convolutions, such as ResNet blocks [3], are popular in large-scale ConvNets; in contrast, depthwise convolutions [28] are popular in mobile platforms due to its lower computational cost and smalle…”
    From CoAtNet · §Related Work
  • MAE2021 · cited 3×
    “Deep learning has witnessed an explosion of architectures of continuously growing capability and capacity Krizhevsky2012; He2016; Vaswani2017.”
    From MAE · §Introduction
  • ViTCAP2021 · cited 2×
    “In xu2021e2e, the image is encoded with ResNet he2016deep and the performance (117.3117.3 CIDEr on COCO xu2021e2e) is still far from the state-of-the-art detector-based approach (129.3129.3 CIDEr with VinVL-base zhang202…”
    From ViTCAP · §Introduction
  • Simple end-to-end captioning2022 · cited 1×, 1 in Method
    “To achieve our goal, we introduce a frustratingly simple, but effective self-ensemble module to combine the output logits of the cross-modal fusion module and the single-modal language decoder with a residual connection…”
    From Simple end-to-end captioning · §. Methodology
  • UniCL2022 · cited 3×
    “When fueled with clean and large-scale human-annotated image-label data, e.g., ImageNet deng2009imagenet, supervised learning can attain decent visual recognition capacities over the given categories krizhevsky2012imagen…”
    From UniCL · §Introduction
  • CoCa2022 · cited 2×, 2 in Method
    “Following a standard encoder-decoder architecture, the image encoder provides latent encoded features (e.g., using a Vision Transformer [39] or ConvNets [40]) and the text decoder learns to maximize the conditional likel…”
    From CoCa · §Approach
  • EVA2022 · cited 2×
    “Notably, the ImageNet-1K validation zero-shot top-1 accuracy is 78.2% without using any of its training set labels, matching the original ResNet-101 resnet.”
    From EVA · §Fly EVA to the Moon
Abstract

Deeper neural networks are more difficult to train. We present a residual learning framework to ease the training of networks that are substantially deeper than those used previously. We explicitly reformulate the layers as learning residual functions with reference to the layer inputs, instead of learning unreferenced functions. We provide comprehensive empirical evidence showing that these residual networks are easier to optimize, and can gain accuracy from considerably increased depth. On the ImageNet dataset we evaluate residual nets with a depth of up to 152 layers---8x deeper than VGG nets but still having lower complexity. An ensemble of these residual nets achieves 3.57% error on the ImageNet test set. This result won the 1st place on the ILSVRC 2015 classification task. We also present analysis on CIFAR-10 with 100 and 1000 layers. The depth of representations is of central importance for many visual recognition tasks. Solely due to our extremely deep representations, we obtain a 28% relative improvement on the COCO object detection dataset. Deep residual nets are foundations of our submissions to ILSVRC & COCO 2015 competitions, where we also won the 1st places on the tasks of ImageNet detection, ImageNet localization, COCO detection, and COCO segmentation.