Measuring Robustness to Natural Distribution Shifts in Image Classification
We study how robust current ImageNet models are to distribution shifts arising from natural variations in datasets. ), which leaves open how robustness on synthetic distribution shift relates to distribution shift arising in real data.
Also cited · not yet reviewed (12)
- ImageNetV22019 · cited 7×, 2 in Method“In order to ensure that the accuracy on the two test sets are comparable, we focus on natural distribution shifts where humans have thoroughly reviewed the test sets to include only correctly labeled images [76, 68, 39,…”From this paper · §Experimental setup
- ImageNet-trained CNNs are biased towards2018 · cited 4×, 2 in Method“We use a stylized version of the ImageNet test set [44, 34].”From this paper · §Experimental setup
- Billion-scale semi-supervised2019 · cited 2×, 2 in Method“This subset includes models trained on (i) Facebook’s collection of 1 billion Instagram images [56, 104], (ii) the YFCC 100 million dataset [104], (iii) Google’s JFT 300 million dataset [82, 102], (iv) a subset of OpenIm…”From this paper · §Experimental setup
- ImageNet-C2019 · cited 5×, 1 in Method“We include all corruptions from [38], as well as some corruptions from [33].”From this paper · §Experimental setup
Show 8 more
- Noisy Student2019 · cited 5×, 1 in Method“This subset includes models trained on (i) Facebook’s collection of 1 billion Instagram images [56, 104], (ii) the YFCC 100 million dataset [104], (iii) Google’s JFT 300 million dataset [82, 102], (iv) a subset of OpenIm…”From this paper · §Experimental setup
- JFT-300M (unreasonable effectiveness)2017 · cited 2×, 1 in Method“This subset includes models trained on (i) Facebook’s collection of 1 billion Instagram images [56, 104], (ii) the YFCC 100 million dataset [104], (iii) Google’s JFT 300 million dataset [82, 102], (iv) a subset of OpenIm…”From this paper · §Experimental setup
- Instagram hashtag pre-training2018 · cited 2×, 1 in Method“This subset includes models trained on (i) Facebook’s collection of 1 billion Instagram images [56, 104], (ii) the YFCC 100 million dataset [104], (iii) Google’s JFT 300 million dataset [82, 102], (iv) a subset of OpenIm…”From this paper · §Experimental setup
- VGG2014 · cited 1×, 1 in Method“This category includes 78 models with architectures ranging from AlexNet to EfficietNet, e.g., [50, 78, 37, 85, 88].”From this paper · §Experimental setup
- GoogLeNet (Inception)2014 · cited 1×, 1 in Method“This category includes 78 models with architectures ranging from AlexNet to EfficietNet, e.g., [50, 78, 37, 85, 88].”From this paper · §Experimental setup
- ResNet2015 · cited 1×, 1 in Method“This category includes 78 models with architectures ranging from AlexNet to EfficietNet, e.g., [50, 78, 37, 85, 88].”From this paper · §Experimental setup
- EfficientNet2019 · cited 1×, 1 in Method“This category includes 78 models with architectures ranging from AlexNet to EfficietNet, e.g., [50, 78, 37, 85, 88].”From this paper · §Experimental setup
- BiT2019 · cited 1×, 1 in Method“This subset includes models trained on (i) Facebook’s collection of 1 billion Instagram images [56, 104], (ii) the YFCC 100 million dataset [104], (iii) Google’s JFT 300 million dataset [82, 102], (iv) a subset of OpenIm…”From this paper · §Experimental setup
Led to
- CLIP2021 · cited 8דEncouragingly however, Taori et al. 2020 find that accuracy under distribution shift increases predictably with ImageNet accuracy and is well modeled as a linear function of logit-transformed accuracy.”From CLIP · §Experiments
Abstract
We study how robust current ImageNet models are to distribution shifts arising from natural variations in datasets. Most research on robustness focuses on synthetic image perturbations (noise, simulated weather artifacts, adversarial examples, etc.), which leaves open how robustness on synthetic distribution shift relates to distribution shift arising in real data. Informed by an evaluation of 204 ImageNet models in 213 different test conditions, we find that there is often little to no transfer of robustness from current synthetic to natural distribution shift. Moreover, most current techniques provide no robustness to the natural distribution shifts in our testbed. The main exception is training on larger and more diverse datasets, which in multiple cases increases robustness, but is still far from closing the performance gaps. Our results indicate that distribution shifts arising in real data are currently an open research problem. We provide our testbed and data as a resource for future work at https://modestyachts.github.io/imagenet-testbed/ .