Paper Lineage
Esc
AnalysisMay 2018arXiv 1805.08974cs.CV

Do Better ImageNet Models Transfer Better?

Simon Kornblith, Jonathon Shlens, Quoc V. Le

Transfer learning is a cornerstone of computer vision, yet little work has been done to evaluate the relationship between architecture and transfer. An implicit hypothesis in modern computer vision research is that models that perform better on ImageNet necessarily perform better on other vision tasks.

From the abstract

Built on

6 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (6)

  • DeCAF2013 · cited 5×
    “Network architectures measured against this dataset have fueled much progress in computer vision research across a broad array of problems, including transferring to new datasets donahue2014decaf; razavian2014cnn, object…”
    From this paper · §Introduction
  • “Network architectures measured against this dataset have fueled much progress in computer vision research across a broad array of problems, including transferring to new datasets donahue2014decaf; razavian2014cnn, object…”
    From this paper · §Introduction
  • “A substantial body of existing research indicates that, in image tasks, fine-tuning typically achieves higher accuracy than classification based on fixed features, especially for larger datasets or datasets with a larger…”
    From this paper · §Related work
  • GoogLeNet (Inception)2014 · cited 3×
    “Although auxiliary classifier heads were initially proposed to alleviate issues related to vanishing gradients lee2015deeply; szegedy2015going, Szegedy et al. szegedy2016rethinking instead suggest that they also act as r…”
    From this paper · §Results
Show 2 more
  • Inception v32015 · cited 3×
    “Although auxiliary classifier heads were initially proposed to alleviate issues related to vanishing gradients lee2015deeply; szegedy2015going, Szegedy et al. szegedy2016rethinking instead suggest that they also act as r…”
    From this paper · §Results
  • MobileNets2017 · cited 2×
    “Chatfield14; simonyan2014very; huang2016; howard2017mobilenets; he2017mask), they have never been systematically explored across network architectures.”
    From this paper · §Introduction

Led to

  • Domain adaptive transfer2018 · cited 2×
    “Kornblith2018 who also found that Inception v3 did slightly better than NasNet-A Zoph2017.”
    From Domain adaptive transfer · §Experiments
  • EfficientNet2019 · cited 2×
    “While these models are mainly designed for ImageNet, recent studies have shown better ImageNet models also perform better across a variety of transfer learning datasets (Kornblith et al. 2019), and other computer vision…”
    From EfficientNet · §Related Work
  • SimCLR2020 · cited 2×
    “Following Kornblith et al. 2019, we perform hyperparameter tuning for each model-dataset combination and select the best hyperparameters on a validation set.”
    From SimCLR · §Comparison with State-of-the-art
  • BYOL2020 · cited 3×
    “We first evaluate BYOL’s representation by training a linear classifier on top of the frozen representation, following the procedure described in [48, 74, 41, 10, 8], and section C.1; we report top-11 and top-55 accuraci…”
    From BYOL · §Experimental evaluation
  • CLIP2021 · cited 6×
    “This increases flexibility, and prior work has convincingly demonstrated that fine-tuning outperforms linear classification on most image classification datasets (Kornblith et al. 2019; Zhai et al. 2019).”
    From CLIP · §Experiments
  • LiT2021 · cited 2×
    “Transfer learning transfer_learning_survey has been a successful paradigm in computer vision imagenet_transfer_better; bit; instagram_resnext.”
    From LiT · §Introduction
  • Florence2021 · cited 2×
    “We evaluate our Florence model on the ImageNet-1K dataset and 11 downstream datasets from the well-studied evaluation suit introduced by (Kornblith et al. 2019).”
    From Florence · §Experiments
Abstract

Transfer learning is a cornerstone of computer vision, yet little work has been done to evaluate the relationship between architecture and transfer. An implicit hypothesis in modern computer vision research is that models that perform better on ImageNet necessarily perform better on other vision tasks. However, this hypothesis has never been systematically tested. Here, we compare the performance of 16 classification networks on 12 image classification datasets. We find that, when networks are used as fixed feature extractors or fine-tuned, there is a strong correlation between ImageNet accuracy and transfer accuracy ($r = 0.99$ and $0.96$, respectively). In the former setting, we find that this relationship is very sensitive to the way in which networks are trained on ImageNet; many common forms of regularization slightly improve ImageNet accuracy but yield penultimate layer features that are much worse for transfer learning. Additionally, we find that, on two small fine-grained image classification datasets, pretraining on ImageNet provides minimal benefits, indicating the learned features from ImageNet do not transfer well to fine-grained tasks. Together, our results show that ImageNet architectures generalize well across datasets, but ImageNet features are less general than previously suggested.