Paper Lineage
Esc
MethodMar 2014arXiv 1403.6382cs.CV

CNN Features off-the-shelf: an Astounding Baseline for Recognition

Ali Sharif Razavian, Hossein Azizpour, Josephine Sullivan, Stefan Carlsson

Recent results indicate that the generic descriptors extracted from the convolutional neural networks are very powerful. This paper adds to the mounting evidence that this is indeed the case.

From the abstract

Built on

2 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (2)

  • OverFeat2013 · cited 3×, 2 in Method
    “In this work we use the publicly available trained CNN called OverFeat Sermanet13.”
    From this paper · §Background and Outline
  • DeCAF2013 · cited 4×
    “The experiments confirm and extend the results reported in Donahue14.”
    From this paper · §Conclusion

Led to

  • “Recent studies have shown that using visual features extracted from convolutional networks trained on large object recognition datasets krizhevsky12; simonyan15; szegedy15 can lead to state-of-the-art results on many vis…”
    From Weakly supervised visual features (Jouli · §Introduction
  • “Network architectures measured against this dataset have fueled much progress in computer vision research across a broad array of problems, including transferring to new datasets donahue2014decaf; razavian2014cnn, object…”
    From Do better ImageNet models transfer bette · §Introduction
  • Domain adaptive transfer2018 · cited 2×
    “The success of applying convolution neural networks to the ImageNet classification problem Krizhevsky2012 led to the finding that the features learned by a convolutional neural network perform well on a variety of image…”
    From Domain adaptive transfer · §Related Work
Abstract

Recent results indicate that the generic descriptors extracted from the convolutional neural networks are very powerful. This paper adds to the mounting evidence that this is indeed the case. We report on a series of experiments conducted for different recognition tasks using the publicly available code and model of the \overfeat network which was trained to perform object classification on ILSVRC13. We use features extracted from the \overfeat network as a generic image representation to tackle the diverse range of recognition tasks of object image classification, scene recognition, fine grained recognition, attribute detection and image retrieval applied to a diverse set of datasets. We selected these tasks and datasets as they gradually move further away from the original task and data the \overfeat network was trained to solve. Astonishingly, we report consistent superior results compared to the highly tuned state-of-the-art systems in all the visual classification tasks on various datasets. For instance retrieval it consistently outperforms low memory footprint methods except for sculptures dataset. The results are achieved using a linear SVM classifier (or $L2$ distance in case of retrieval) applied to a feature representation of size 4096 extracted from a layer in the net. The representations are further modified using simple augmentation techniques e.g. jittering. The results strongly suggest that features obtained from deep learning with convolutional nets should be the primary candidate in most visual recognition tasks.