Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
The success of deep learning in vision can be attributed to: (a) models with high capacity; (b) increased computational power; and (c) availability of large-scale labeled data. Since 2012, there have been significant advances in representation capabilities of the models and computational capabilities of GPUs.
Also cited · not yet reviewed (6)
- ResNet2015 · cited 6×, 3 in Method“Our baseline ResNet-101 performs 1% better than the open-sourced ResNet-101 checkpoint from the authors of [16], using the same evaluation protocol.”From this paper · §Training and Evaluation Framework
- Faster R-CNN2015 · cited 4×, 1 in Method“We use the Faster RCNN framework [33] for its state-of-the-art performance.”From this paper · §Training and Evaluation Framework
- BatchNorm2015 · cited 1×, 1 in Method“We set weight decay to 10−410^{-4} and use batch normalization [20] after all the convolutional layers.”From this paper · §Training and Evaluation Framework
- Weakly supervised visual features (Jouli2015 · cited 5דIn order to overcome the bottleneck, there have been recent efforts on visual representation learning using web-supervision [5, 6, 9, 21, 2, 23, 27, 24] or unsupervised [34, 10, 11, 43, 31, 32, 42] paradigms.”From this paper · §Related Work
Show 2 more
- MS COCO2014 · cited 2דWe evaluate on the two most popular datasets: COCO [26] and PASCAL VOC [13].”From this paper · §Experiments
- Knowledge Distillation2015 · cited 2דJFT-300M is a follow up version of the dataset introduced by [17, 7].”From this paper · §The JFT-300M Dataset
Led to
- Instagram hashtag pre-training2018 · cited 8×, 3 in Method“For instance, ∼5%\raise 0.73193pt\hbox{$\scriptstyle\sim$}5\% of the images in the val-CUB-6k-200 set [21] also appear in train-IN-1M-1k, and 1.78%1.78\% of images in val-IN-50k-1k set are in the JFT-300M training set [1…”From Instagram hashtag pre-training · §Scaling up Supervised Pretraining
- Domain adaptive transfer2018 · cited 2×, 2 in Method“We use the JFT Sun2017 and ImageNet Russakovsky2015 datasets as our source pre-training data and consider a range of target datasets for fine-tuning (Section 3.2).”From Domain adaptive transfer · §Transfer learning setup
- Natural distribution shift robustness2020 · cited 2×, 1 in Method“This subset includes models trained on (i) Facebook’s collection of 1 billion Instagram images [56, 104], (ii) the YFCC 100 million dataset [104], (iii) Google’s JFT 300 million dataset [82, 102], (iv) a subset of OpenIm…”From Natural distribution shift robustness · §Experimental setup
- ViT2020 · cited 2דMoreover, Sun et al. 2017 study how CNN performance scales with dataset size, and Kolesnikov et al. 2020; Djolonga et al. 2020 perform an empirical exploration of CNN transfer learning from large scale datasets such as I…”From ViT · §Related Work
- DALL·E2021 · cited 1×, 1 in Method“To scale up to 1212-billion parameters, we created a dataset of a similar scale to JFT-300M (Sun et al. 2017) by collecting 250 million text-images pairs from the internet.”From DALL·E · §Method
- Scaling ViTs (ViT-G)2021 · cited 1×, 1 in Method“For this study, we use the proprietary JFT-3B dataset, a larger version of the JFT-300M dataset used in many previous works on large-scale computer vision models sun2017unreasonable; kolesnikov2019big; dosovitskiy2020.”From Scaling ViTs (ViT-G) · §Method details
- LiT2021 · cited 1×, 1 in Method“Some dataset choices for learning powerful image embeddings are ImageNet-21k imagenet, JFT-300M unreasonable_effectiveness_of_data.”From LiT · §Methods
Abstract
The success of deep learning in vision can be attributed to: (a) models with high capacity; (b) increased computational power; and (c) availability of large-scale labeled data. Since 2012, there have been significant advances in representation capabilities of the models and computational capabilities of GPUs. But the size of the biggest dataset has surprisingly remained constant. What will happen if we increase the dataset size by 10x or 100x? This paper takes a step towards clearing the clouds of mystery surrounding the relationship between `enormous data' and visual deep learning. By exploiting the JFT-300M dataset which has more than 375M noisy labels for 300M images, we investigate how the performance of current vision tasks would change if this data was used for representation learning. Our paper delivers some surprising (and some expected) findings. First, we find that the performance on vision tasks increases logarithmically based on volume of training data size. Second, we show that representation learning (or pre-training) still holds a lot of promise. One can improve performance on many vision tasks by just training a better base model. Finally, as expected, we present new state-of-the-art results for different vision tasks including image classification, object detection, semantic segmentation and human pose estimation. Our sincere hope is that this inspires vision community to not undervalue the data and develop collective efforts in building larger datasets.