Billion-scale semi-supervised learning for image classification
This paper presents a study of semi-supervised learning with large convolutional networks. We propose a pipeline, based on a teacher/student paradigm, that leverages a large collection of unlabelled images (up to 1 billion).
Also cited · not yet reviewed (5)
- Instagram hashtag pre-training2018 · cited 8דIG-1B-Targeted: Following [27], we collected a dataset of 1B public images with associated hashtags from a social media website.”From this paper · §Image classification: experiments & analysis
- ResNet2015 · cited 4דModels: For student and teacher models, we use residual networks [16], ResNet-d with dd = {18,50}\{18,50\} and residual networks with group convolutions [43], ResNeXt-101 32XCd with 101101 layers and group widths CC = {4…”From this paper · §Image classification: experiments & analysis
- ResNeXt2016 · cited 3דModels: For student and teacher models, we use residual networks [16], ResNet-d with dd = {18,50}\{18,50\} and residual networks with group convolutions [43], ResNeXt-101 32XCd with 101101 layers and group widths CC = {4…”From this paper · §Image classification: experiments & analysis
- BatchNorm2015 · cited 2דEach GPU processes 2424 images at a time and apply batch normalization [22] to all convolutional layers on each GPU.”From this paper · §Image classification: experiments & analysis
Show 1 more
- Goyal large-batch SGD2017 · cited 2דWe set the learning rate following the linear scaling procedure proposed in [13] with a warm-up and overall minibatch size of 64×24=153664\times 24=1536.”From this paper · §Image classification: experiments & analysis
Led to
- Noisy Student2019 · cited 3דOther Techniques – Noisy Student Training also works better with an additional trick: data filtering and balancing, similar to uda; billion_large_scale.”From Noisy Student · §Noisy Student Training
- BiT2019 · cited 4דRather than pre-train generic representations, recent works have shown strong performance by training task-specific representations [63, 38, 61].”From BiT · §Related Work
- Natural distribution shift robustness2020 · cited 2×, 2 in Method“This subset includes models trained on (i) Facebook’s collection of 1 billion Instagram images [56, 104], (ii) the YFCC 100 million dataset [104], (iii) Google’s JFT 300 million dataset [82, 102], (iv) a subset of OpenIm…”From Natural distribution shift robustness · §Experimental setup
Abstract
This paper presents a study of semi-supervised learning with large convolutional networks. We propose a pipeline, based on a teacher/student paradigm, that leverages a large collection of unlabelled images (up to 1 billion). Our main goal is to improve the performance for a given target architecture, like ResNet-50 or ResNext. We provide an extensive analysis of the success factors of our approach, which leads us to formulate some recommendations to produce high-accuracy models for image classification with semi-supervised learning. As a result, our approach brings important gains to standard architectures for image, video and fine-grained classification. For instance, by leveraging one billion unlabelled images, our learned vanilla ResNet-50 achieves 81.2% top-1 accuracy on the ImageNet benchmark.