Benchmarking Neural Network Robustness to Common Corruptions and Perturbations
In this paper we establish rigorous benchmarks for image classifier robustness. Our first benchmark, ImageNet-C, standardizes and expands the corruption robustness topic, while showing which classifiers are preferable in safety-critical applications.
Also cited · not yet reviewed (1)
- ResNeXt2016 · cited 2דUnlike current deep learning classifiers (Krizhevsky et al. 2012; He et al. 2015; Xie et al. 2016), the human vision system is not fooled by small changes in query images.”From this paper · §Introduction
Led to
- Noisy Student2019 · cited 4דNot only our method improves standard ImageNet accuracy, it also improves classification robustness on much harder test sets by large margins: ImageNet-A hendrycks2019natural top-1 accuracy from 61.0% to 83.7%, ImageNet-…”From Noisy Student · §Introduction
- Natural distribution shift robustness2020 · cited 5×, 1 in Method“We include all corruptions from [38], as well as some corruptions from [33].”From Natural distribution shift robustness · §Experimental setup
Abstract
In this paper we establish rigorous benchmarks for image classifier robustness. Our first benchmark, ImageNet-C, standardizes and expands the corruption robustness topic, while showing which classifiers are preferable in safety-critical applications. Then we propose a new dataset called ImageNet-P which enables researchers to benchmark a classifier's robustness to common perturbations. Unlike recent robustness research, this benchmark evaluates performance on common corruptions and perturbations not worst-case adversarial perturbations. We find that there are negligible changes in relative corruption robustness from AlexNet classifiers to ResNet classifiers. Afterward we discover ways to enhance corruption and perturbation robustness. We even find that a bypassed adversarial defense provides substantial common perturbation robustness. Together our benchmarks may aid future work toward networks that robustly generalize.