Do Better ImageNet Models Transfer Better?
Transfer learning is a cornerstone of computer vision, yet little work has been done to evaluate the relationship between architecture and transfer. An implicit hypothesis in modern computer vision research is that models that perform better on ImageNet necessarily perform better on other vision tasks.
Also cited · not yet reviewed (6)
- DeCAF2013 · cited 5דNetwork architectures measured against this dataset have fueled much progress in computer vision research across a broad array of problems, including transferring to new datasets donahue2014decaf; razavian2014cnn, object…”From this paper · §Introduction
- CNN features off-the-shelf2014 · cited 5דNetwork architectures measured against this dataset have fueled much progress in computer vision research across a broad array of problems, including transferring to new datasets donahue2014decaf; razavian2014cnn, object…”From this paper · §Introduction
- Return of the Devil (CNN details)2014 · cited 5דA substantial body of existing research indicates that, in image tasks, fine-tuning typically achieves higher accuracy than classification based on fixed features, especially for larger datasets or datasets with a larger…”From this paper · §Related work
- GoogLeNet (Inception)2014 · cited 3דAlthough auxiliary classifier heads were initially proposed to alleviate issues related to vanishing gradients lee2015deeply; szegedy2015going, Szegedy et al. szegedy2016rethinking instead suggest that they also act as r…”From this paper · §Results
Show 2 more
- Inception v32015 · cited 3דAlthough auxiliary classifier heads were initially proposed to alleviate issues related to vanishing gradients lee2015deeply; szegedy2015going, Szegedy et al. szegedy2016rethinking instead suggest that they also act as r…”From this paper · §Results
- MobileNets2017 · cited 2דChatfield14; simonyan2014very; huang2016; howard2017mobilenets; he2017mask), they have never been systematically explored across network architectures.”From this paper · §Introduction
Led to
- Domain adaptive transfer2018 · cited 2דKornblith2018 who also found that Inception v3 did slightly better than NasNet-A Zoph2017.”From Domain adaptive transfer · §Experiments
- EfficientNet2019 · cited 2דWhile these models are mainly designed for ImageNet, recent studies have shown better ImageNet models also perform better across a variety of transfer learning datasets (Kornblith et al. 2019), and other computer vision…”From EfficientNet · §Related Work
- SimCLR2020 · cited 2דFollowing Kornblith et al. 2019, we perform hyperparameter tuning for each model-dataset combination and select the best hyperparameters on a validation set.”From SimCLR · §Comparison with State-of-the-art
- BYOL2020 · cited 3דWe first evaluate BYOL’s representation by training a linear classifier on top of the frozen representation, following the procedure described in [48, 74, 41, 10, 8], and section C.1; we report top-11 and top-55 accuraci…”From BYOL · §Experimental evaluation
- CLIP2021 · cited 6דThis increases flexibility, and prior work has convincingly demonstrated that fine-tuning outperforms linear classification on most image classification datasets (Kornblith et al. 2019; Zhai et al. 2019).”From CLIP · §Experiments
- LiT2021 · cited 2דTransfer learning transfer_learning_survey has been a successful paradigm in computer vision imagenet_transfer_better; bit; instagram_resnext.”From LiT · §Introduction
- Florence2021 · cited 2דWe evaluate our Florence model on the ImageNet-1K dataset and 11 downstream datasets from the well-studied evaluation suit introduced by (Kornblith et al. 2019).”From Florence · §Experiments
Abstract
Transfer learning is a cornerstone of computer vision, yet little work has been done to evaluate the relationship between architecture and transfer. An implicit hypothesis in modern computer vision research is that models that perform better on ImageNet necessarily perform better on other vision tasks. However, this hypothesis has never been systematically tested. Here, we compare the performance of 16 classification networks on 12 image classification datasets. We find that, when networks are used as fixed feature extractors or fine-tuned, there is a strong correlation between ImageNet accuracy and transfer accuracy ($r = 0.99$ and $0.96$, respectively). In the former setting, we find that this relationship is very sensitive to the way in which networks are trained on ImageNet; many common forms of regularization slightly improve ImageNet accuracy but yield penultimate layer features that are much worse for transfer learning. Additionally, we find that, on two small fine-grained image classification datasets, pretraining on ImageNet provides minimal benefits, indicating the learned features from ImageNet do not transfer well to fine-grained tasks. Together, our results show that ImageNet architectures generalize well across datasets, but ImageNet features are less general than previously suggested.