Exploring the Limits of Weakly Supervised Pretraining
State-of-the-art visual perception models for a wide range of tasks rely on supervised pretraining. ImageNet classification is the de facto pretraining task for these models.
Also cited · not yet reviewed (8)
- Goyal large-batch SGD2017 · cited 4×, 4 in Method“To make training at this scale practical, we adopt a distributed synchronous implementation of stochastic gradient descent with large (8k image) minibatches, following Goyal et al. [19].”From this paper · §Scaling up Supervised Pretraining
- JFT-300M (unreasonable effectiveness)2017 · cited 8×, 3 in Method“For instance, ∼5%\raise 0.73193pt\hbox{$\scriptstyle\sim$}5\% of the images in the val-CUB-6k-200 set [21] also appear in train-IN-1M-1k, and 1.78%1.78\% of images in val-IN-50k-1k set are in the JFT-300M training set [1…”From this paper · §Scaling up Supervised Pretraining
- ResNeXt2016 · cited 3×, 2 in Method“For the 5k set, we use the now standard IN-5k proposed in [15] (6.6M training images).”From this paper · §Scaling up Supervised Pretraining
- Weakly supervised visual features (Jouli2015 · cited 7×, 1 in Method“Motivated by this work, we perform experiments in which we evaluate three different types of data sampling in the Instagram pretraining: (1) a natural sampling in which we sample images and hashtags according to the dist…”From this paper · §Experiments
Show 4 more
- PReLU / He init2015 · cited 1×, 1 in Method“We use a warm-up from 0.1 up to 0.1/256×80640.1/256\times 8064, where 0.1 and 256 are canonical learning rate and minibatch sizes [28].”From this paper · §Scaling up Supervised Pretraining
- BatchNorm2015 · cited 1×, 1 in Method“Each GPU processes 24 images at a time and batch normalization (BN) [27] statistics are computed on these 24 image sets.”From this paper · §Scaling up Supervised Pretraining
- ResNet2015 · cited 1×, 1 in Method“We believe our results will generalize to other architectures [24, 25, 26].”From this paper · §Scaling up Supervised Pretraining
- MS COCO2014 · cited 2דFor example, we observe improvements over the state-of-the-art for image classification and object detection, where we obtain a single-crop, top-1 accuracy of 85.4% on the ImageNet-1k image-classification dataset and 45.…”From this paper · §Introduction
Led to
- Billion-scale semi-supervised2019 · cited 8דIG-1B-Targeted: Following [27], we collected a dataset of 1B public images with associated hashtags from a social media website.”From Billion-scale semi-supervised · §Image classification: experiments & analysis
- T52019 · cited 2דThis is a natural fit for neural networks, which have been shown to exhibit remarkable scalability, i.e. it is often possible to achieve better performance simply by training a larger model on a larger data set (Hestness…”From T5 · §Introduction
- Noisy Student2019 · cited 3דFurther, Noisy Student Training outperforms the state-of-the-art accuracy of 86.4% by FixRes ResNeXt-101 WSL mahajan2018exploring; touvron2019fixing that requires 3.5 Billion Instagram images labeled with tags.”From Noisy Student · §Experiments
- MoCo2019 · cited 2דInstagram-1B (IG-1B): Following Mahajan2018, this is a dataset of ∼\scriptstyle\sim1 billion (940M) public images from Instagram.”From MoCo · §Experiments
- Natural distribution shift robustness2020 · cited 2×, 1 in Method“This subset includes models trained on (i) Facebook’s collection of 1 billion Instagram images [56, 104], (ii) the YFCC 100 million dataset [104], (iii) Google’s JFT 300 million dataset [82, 102], (iv) a subset of OpenIm…”From Natural distribution shift robustness · §Experimental setup
- ViT2020 · cited 2דThe use of additional data sources allows to achieve state-of-the-art results on standard benchmarks (Mahajan et al. 2018; Touvron et al. 2019; Xie et al. 2020).”From ViT · §Related Work
- CLIP2021 · cited 9×, 3 in Method“By comparison, other computer vision systems are trained on up to 3.5 billion Instagram photos (Mahajan et al. 2018).”From CLIP · §Approach
- Scaling ViTs (ViT-G)2021 · cited 2×, 1 in Method“Inspired by instagram, we address this issue by exploring learning-rate schedules that, similar to the warmup phase in the beginning, include a cooldown phase at the end of training, where the learning-rate is linearly a…”From Scaling ViTs (ViT-G) · §Method details
- CoCa2022 · cited 2×, 1 in Method“The classic single-encoder approach pretrains a visual encoder through image classification on a large crowd-sourced image annotation dataset (e.g., ImageNet [9], Instagram [20] or JFT [21]), where the vocabulary of anno…”From CoCa · §Approach
Abstract
State-of-the-art visual perception models for a wide range of tasks rely on supervised pretraining. ImageNet classification is the de facto pretraining task for these models. Yet, ImageNet is now nearly ten years old and is by modern standards "small". Even so, relatively little is known about the behavior of pretraining with datasets that are multiple orders of magnitude larger. The reasons are obvious: such datasets are difficult to collect and annotate. In this paper, we present a unique study of transfer learning with large convolutional networks trained to predict hashtags on billions of social media images. Our experiments demonstrate that training for large-scale hashtag prediction leads to excellent results. We show improvements on several image classification and object detection tasks, and report the highest ImageNet-1k single-crop, top-1 accuracy to date: 85.4% (97.6% top-5). We also perform extensive experiments that provide novel empirical data on the relationship between large-scale pretraining and transfer learning performance.