ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness
Convolutional Neural Networks (CNNs) are commonly thought to recognise objects by learning increasingly complex representations of object shapes. Some recent studies suggest a more important role of image textures.
Also cited · not yet reviewed (3)
- VGG2014 · cited 1×, 1 in Method“The same images were fed to four CNNs pre-trained on standard ImageNet, namely AlexNet (Krizhevsky et al. 2012), GoogLeNet (Szegedy et al. 2015), VGG-16 (Simonyan & Zisserman 2015) and ResNet-50 (He et al. 2015).”From this paper · §Methods
- GoogLeNet (Inception)2014 · cited 1×, 1 in Method“The same images were fed to four CNNs pre-trained on standard ImageNet, namely AlexNet (Krizhevsky et al. 2012), GoogLeNet (Szegedy et al. 2015), VGG-16 (Simonyan & Zisserman 2015) and ResNet-50 (He et al. 2015).”From this paper · §Methods
- ResNet2015 · cited 1×, 1 in Method“The same images were fed to four CNNs pre-trained on standard ImageNet, namely AlexNet (Krizhevsky et al. 2012), GoogLeNet (Szegedy et al. 2015), VGG-16 (Simonyan & Zisserman 2015) and ResNet-50 (He et al. 2015).”From this paper · §Methods
Led to
- Natural distribution shift robustness2020 · cited 4×, 2 in Method“We use a stylized version of the ImageNet test set [44, 34].”From Natural distribution shift robustness · §Experimental setup
- CLIP2021 · cited 2דHowever, research in the subsequent years has repeatedly found that these models still make many simple mistakes (Dodge & Karam 2017; Geirhos et al. 2018; Alcorn et al. 2019), and new benchmarks testing these systems has…”From CLIP · §Experiments
Abstract
Convolutional Neural Networks (CNNs) are commonly thought to recognise objects by learning increasingly complex representations of object shapes. Some recent studies suggest a more important role of image textures. We here put these conflicting hypotheses to a quantitative test by evaluating CNNs and human observers on images with a texture-shape cue conflict. We show that ImageNet-trained CNNs are strongly biased towards recognising textures rather than shapes, which is in stark contrast to human behavioural evidence and reveals fundamentally different classification strategies. We then demonstrate that the same standard architecture (ResNet-50) that learns a texture-based representation on ImageNet is able to learn a shape-based representation instead when trained on "Stylized-ImageNet", a stylized version of ImageNet. This provides a much better fit for human behavioural performance in our well-controlled psychophysical lab setting (nine experiments totalling 48,560 psychophysical trials across 97 observers) and comes with a number of unexpected emergent benefits such as improved object detection performance and previously unseen robustness towards a wide range of image distortions, highlighting advantages of a shape-based representation.