Paper Lineage
Esc
MethodNov 2019arXiv 1911.04252cs.LG

Self-training with Noisy Student improves ImageNet classification

Qizhe Xie, Minh-Thang Luong, Eduard Hovy, Quoc V. Le

We present Noisy Student Training, a semi-supervised learning approach that works well even when labeled data is abundant. 5B weakly labeled Instagram images.

From the abstract

Built on

8 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (8)

  • EfficientNet2019 · cited 6×
    “Deep learning has shown remarkable successes in image recognition in recent years krizhevsky2012imagenet; szegedy2015going; simonyan2014very; he2016deep; tan2019efficientnet.”
    From this paper · §Introduction
  • Knowledge Distillation2015 · cited 5×
    “The algorithm is an improved version of self-training, a method in semi-supervised learning (e.g., scudder1965probability; yarowsky1995unsupervised), and distillation hinton2015distilling.”
    From this paper · §Noisy Student Training
  • ImageNet-C2019 · cited 4×
    “Not only our method improves standard ImageNet accuracy, it also improves classification robustness on much harder test sets by large margins: ImageNet-A hendrycks2019natural top-1 accuracy from 61.0% to 83.7%, ImageNet-…”
    From this paper · §Introduction
  • “Further, Noisy Student Training outperforms the state-of-the-art accuracy of 86.4% by FixRes ResNeXt-101 WSL mahajan2018exploring; touvron2019fixing that requires 3.5 Billion Instagram images labeled with tags.”
    From this paper · §Experiments
Show 4 more
  • ImageNetV22019 · cited 3×
    “We conduct experiments on ImageNet 2012 ILSVRC challenge prediction task since it has been considered one of the most heavily benchmarked datasets in computer vision and that improvements on ImageNet transfer to other da…”
    From this paper · §Experiments
  • “Other Techniques – Noisy Student Training also works better with an additional trick: data filtering and balancing, similar to uda; billion_large_scale.”
    From this paper · §Noisy Student Training
  • GoogLeNet (Inception)2014 · cited 2×
    “Deep learning has shown remarkable successes in image recognition in recent years krizhevsky2012imagenet; szegedy2015going; simonyan2014very; he2016deep; tan2019efficientnet.”
    From this paper · §Introduction
  • ResNet2015 · cited 2×
    “Deep learning has shown remarkable successes in image recognition in recent years krizhevsky2012imagenet; szegedy2015going; simonyan2014very; he2016deep; tan2019efficientnet.”
    From this paper · §Introduction

Led to

  • BiT2019 · cited 4×
    “Rather than pre-train generic representations, recent works have shown strong performance by training task-specific representations [63, 38, 61].”
    From BiT · §Related Work
  • Natural distribution shift robustness2020 · cited 5×, 1 in Method
    “This subset includes models trained on (i) Facebook’s collection of 1 billion Instagram images [56, 104], (ii) the YFCC 100 million dataset [104], (iii) Google’s JFT 300 million dataset [82, 102], (iv) a subset of OpenIm…”
    From Natural distribution shift robustness · §Experimental setup
  • ViT2020 · cited 3×
    “The use of additional data sources allows to achieve state-of-the-art results on standard benchmarks (Mahajan et al. 2018; Touvron et al. 2019; Xie et al. 2020).”
    From ViT · §Related Work
  • CLIP2021 · cited 3×, 1 in Method
    “Mahajan et al. 2018 required 19 GPU years to train their ResNeXt101-32x48d and Xie et al. 2020 required 33 TPUv3 core-years to train their Noisy Student EfficientNet-L2.”
    From CLIP · §Approach
  • Scaling ViTs (ViT-G)2021 · cited 2×
    “On ImageNet-v2, ViT-G/14 improves 3% over the Noisy Student model xie2019selftraining based on EfficientNet-L2.”
    From Scaling ViTs (ViT-G) · §Core Results
  • Florence2021 · cited 3×
    “Linear probe as another main metric for evaluating representation quality has been used in most recent studies, including self-supervised learning (Chen et al. 2020b; Chen et al. 2020c), self-training with noisy student…”
    From Florence · §Experiments
Abstract

We present Noisy Student Training, a semi-supervised learning approach that works well even when labeled data is abundant. Noisy Student Training achieves 88.4% top-1 accuracy on ImageNet, which is 2.0% better than the state-of-the-art model that requires 3.5B weakly labeled Instagram images. On robustness test sets, it improves ImageNet-A top-1 accuracy from 61.0% to 83.7%, reduces ImageNet-C mean corruption error from 45.7 to 28.3, and reduces ImageNet-P mean flip rate from 27.8 to 12.2. Noisy Student Training extends the idea of self-training and distillation with the use of equal-or-larger student models and noise added to the student during learning. On ImageNet, we first train an EfficientNet model on labeled images and use it as a teacher to generate pseudo labels for 300M unlabeled images. We then train a larger EfficientNet as a student model on the combination of labeled and pseudo labeled images. We iterate this process by putting back the student as the teacher. During the learning of the student, we inject noise such as dropout, stochastic depth, and data augmentation via RandAugment to the student so that the student generalizes better than the teacher. Models are available at https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet. Code is available at https://github.com/google-research/noisystudent.