Self-training with Noisy Student improves ImageNet classification
We present Noisy Student Training, a semi-supervised learning approach that works well even when labeled data is abundant. 5B weakly labeled Instagram images.
Also cited · not yet reviewed (8)
- EfficientNet2019 · cited 6דDeep learning has shown remarkable successes in image recognition in recent years krizhevsky2012imagenet; szegedy2015going; simonyan2014very; he2016deep; tan2019efficientnet.”From this paper · §Introduction
- Knowledge Distillation2015 · cited 5דThe algorithm is an improved version of self-training, a method in semi-supervised learning (e.g., scudder1965probability; yarowsky1995unsupervised), and distillation hinton2015distilling.”From this paper · §Noisy Student Training
- ImageNet-C2019 · cited 4דNot only our method improves standard ImageNet accuracy, it also improves classification robustness on much harder test sets by large margins: ImageNet-A hendrycks2019natural top-1 accuracy from 61.0% to 83.7%, ImageNet-…”From this paper · §Introduction
- Instagram hashtag pre-training2018 · cited 3דFurther, Noisy Student Training outperforms the state-of-the-art accuracy of 86.4% by FixRes ResNeXt-101 WSL mahajan2018exploring; touvron2019fixing that requires 3.5 Billion Instagram images labeled with tags.”From this paper · §Experiments
Show 4 more
- ImageNetV22019 · cited 3דWe conduct experiments on ImageNet 2012 ILSVRC challenge prediction task since it has been considered one of the most heavily benchmarked datasets in computer vision and that improvements on ImageNet transfer to other da…”From this paper · §Experiments
- Billion-scale semi-supervised2019 · cited 3דOther Techniques – Noisy Student Training also works better with an additional trick: data filtering and balancing, similar to uda; billion_large_scale.”From this paper · §Noisy Student Training
- GoogLeNet (Inception)2014 · cited 2דDeep learning has shown remarkable successes in image recognition in recent years krizhevsky2012imagenet; szegedy2015going; simonyan2014very; he2016deep; tan2019efficientnet.”From this paper · §Introduction
- ResNet2015 · cited 2דDeep learning has shown remarkable successes in image recognition in recent years krizhevsky2012imagenet; szegedy2015going; simonyan2014very; he2016deep; tan2019efficientnet.”From this paper · §Introduction
Led to
- BiT2019 · cited 4דRather than pre-train generic representations, recent works have shown strong performance by training task-specific representations [63, 38, 61].”From BiT · §Related Work
- Natural distribution shift robustness2020 · cited 5×, 1 in Method“This subset includes models trained on (i) Facebook’s collection of 1 billion Instagram images [56, 104], (ii) the YFCC 100 million dataset [104], (iii) Google’s JFT 300 million dataset [82, 102], (iv) a subset of OpenIm…”From Natural distribution shift robustness · §Experimental setup
- ViT2020 · cited 3דThe use of additional data sources allows to achieve state-of-the-art results on standard benchmarks (Mahajan et al. 2018; Touvron et al. 2019; Xie et al. 2020).”From ViT · §Related Work
- CLIP2021 · cited 3×, 1 in Method“Mahajan et al. 2018 required 19 GPU years to train their ResNeXt101-32x48d and Xie et al. 2020 required 33 TPUv3 core-years to train their Noisy Student EfficientNet-L2.”From CLIP · §Approach
- Scaling ViTs (ViT-G)2021 · cited 2דOn ImageNet-v2, ViT-G/14 improves 3% over the Noisy Student model xie2019selftraining based on EfficientNet-L2.”From Scaling ViTs (ViT-G) · §Core Results
- Florence2021 · cited 3דLinear probe as another main metric for evaluating representation quality has been used in most recent studies, including self-supervised learning (Chen et al. 2020b; Chen et al. 2020c), self-training with noisy student…”From Florence · §Experiments
Abstract
We present Noisy Student Training, a semi-supervised learning approach that works well even when labeled data is abundant. Noisy Student Training achieves 88.4% top-1 accuracy on ImageNet, which is 2.0% better than the state-of-the-art model that requires 3.5B weakly labeled Instagram images. On robustness test sets, it improves ImageNet-A top-1 accuracy from 61.0% to 83.7%, reduces ImageNet-C mean corruption error from 45.7 to 28.3, and reduces ImageNet-P mean flip rate from 27.8 to 12.2. Noisy Student Training extends the idea of self-training and distillation with the use of equal-or-larger student models and noise added to the student during learning. On ImageNet, we first train an EfficientNet model on labeled images and use it as a teacher to generate pseudo labels for 300M unlabeled images. We then train a larger EfficientNet as a student model on the combination of labeled and pseudo labeled images. We iterate this process by putting back the student as the teacher. During the learning of the student, we inject noise such as dropout, stochastic depth, and data augmentation via RandAugment to the student so that the student generalizes better than the teacher. Models are available at https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet. Code is available at https://github.com/google-research/noisystudent.