Paper Lineage
Esc
BenchmarkFeb 2019arXiv 1902.10811cs.CV

Do ImageNet Classifiers Generalize to ImageNet?

Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, Vaishaal Shankar

We build new test sets for the CIFAR-10 and ImageNet datasets. Both benchmarks have been the focus of intense research for almost a decade, raising the danger of overfitting to excessively re-used test sets.

From the abstract

Built on

3 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (3)

  • VGG2014 · cited 2×
    “The models include the seminal AlexNet [36], widely used convolutional networks [49, 21, 27, 52], and the state-of-the-art [8, 39].”
    From this paper · §Summary of Our Experiments
  • PReLU / He init2015 · cited 2×
    “State-of-the-art models now surpass human-level accuracy by some measure [20, 48].”
    From this paper · §Summary of Our Experiments
  • ResNet2015 · cited 2×
    “The models include the seminal AlexNet [36], widely used convolutional networks [49, 21, 27, 52], and the state-of-the-art [8, 39].”
    From this paper · §Summary of Our Experiments

Led to

  • Noisy Student2019 · cited 3×
    “We conduct experiments on ImageNet 2012 ILSVRC challenge prediction task since it has been considered one of the most heavily benchmarked datasets in computer vision and that improvements on ImageNet transfer to other da…”
    From Noisy Student · §Experiments
  • Natural distribution shift robustness2020 · cited 7×, 2 in Method
    “In order to ensure that the accuracy on the two test sets are comparable, we focus on natural distribution shifts where humans have thoroughly reviewed the test sets to include only correctly labeled images [76, 68, 39,…”
    From Natural distribution shift robustness · §Experimental setup
  • ALIGN2021 · cited 1×, 1 in Method
    “We first apply zero-shot transfer of ALIGN to visual classification tasks on ImageNet ILSVRC-2012 benchmark (Deng et al. 2009) and its variants including ImageNet-R(endition) (Hendrycks et al. 2020) (non-natural images s…”
    From ALIGN · §Pre-training and Task Transfer
  • CLIP2021 · cited 2×
    “However, research in the subsequent years has repeatedly found that these models still make many simple mistakes (Dodge & Karam 2017; Geirhos et al. 2018; Alcorn et al. 2019), and new benchmarks testing these systems has…”
    From CLIP · §Experiments
  • Scaling ViTs (ViT-G)2021 · cited 2×
    “In addition to ImageNet fine-tuning and linear 10-shot results on the public validation set, we also report results of the ImageNet fine-tuned model on the ImageNet-v2 test set recht2019imagenet as an indicator of robust…”
    From Scaling ViTs (ViT-G) · §Core Results
  • EVA2022 · cited 2×
    “We also evaluate the robustness & generalization capability of EVA along with our training settings & hyper-parameters using ImageNet-V2 matched frequency (IN-V2) inv2, ImageNet-ReaL (IN-ReaL) inreal, ImageNet-Adversaria…”
    From EVA · §Fly EVA to the Moon
Abstract

We build new test sets for the CIFAR-10 and ImageNet datasets. Both benchmarks have been the focus of intense research for almost a decade, raising the danger of overfitting to excessively re-used test sets. By closely following the original dataset creation processes, we test to what extent current classification models generalize to new data. We evaluate a broad range of models and find accuracy drops of 3% - 15% on CIFAR-10 and 11% - 14% on ImageNet. However, accuracy gains on the original test sets translate to larger gains on the new test sets. Our results suggest that the accuracy drops are not caused by adaptivity, but by the models' inability to generalize to slightly "harder" images than those found in the original test sets.