Paper Lineage
Esc
MethodMay 2019arXiv 1905.09272cs.CV

Data-Efficient Image Recognition with Contrastive Predictive Coding

Olivier J. Hénaff, Aravind Srinivas, Jeffrey De Fauw and 4 others

Human observers can learn to recognize new categories of images from a handful of examples, yet doing so with artificial ones remains an open challenge. We hypothesize that data-efficient recognition is enabled by representations which make the variability in natural signals more predictable.

From the abstract

Built on

4 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (4)

  • CPC2018 · cited 8×, 3 in Method
    “This loss is called InfoNCE as it is inspired by Noise-Contrastive Estimation (Gutmann & Hyvärinen 2010; Mnih & Kavukcuoglu 2013) and has been shown to maximize the mutual information between 𝒄i,j{\bm{c}}_{i,j} and 𝒛i+…”
    From this paper · §Experimental Setup
  • Faster R-CNN2015 · cited 2×, 1 in Method
    “As such 𝔻l{\mathbb{D}}_{l} is the entire PASCAL VOC 2007 dataset (comprised of 5011 labeled images); hψh_{\psi} and ℒSupL_{\textrm{Sup}} are the Faster-RCNN architecture and loss (Ren et al. 2015).”
    From this paper · §Experimental Setup
  • ResNet2015 · cited 1×, 1 in Method
    “For this task, the classifier hψh_{\psi} is an arbitrary deep neural network (we use an 11-block ResNet architecture (He et al. 2016a) with 4096-dimensional feature maps and 1024-dimensional bottleneck layers).”
    From this paper · §Experimental Setup
  • AMDIM2019 · cited 2×
    “Augmented Multiscale Deep InfoMax (AMDIM, Bachman et al. 2019) is most similar to CPC in that it makes predictions across space, but differs in that it also predicts representations across layers in the model.”
    From this paper · §Related Work

Led to

  • MoCo2019 · cited 5×, 2 in Method
    “Several recent studies Wu2018a; Oord2018; Hjelm2019; Zhuang2019; Henaff2019; Tian2019; Bachman2019 present promising results on unsupervised visual representation learning using approaches related to the contrastive loss…”
    From MoCo · §Introduction
  • SimCLR2020 · cited 10×, 1 in Method
    “Not only does SimCLR outperform previous work (Figure 1), but it is also simpler, requiring neither specialized architectures (Bachman et al. 2019; Hénaff et al. 2019) nor a memory bank (Wu et al. 2018; Tian et al. 2019;…”
    From SimCLR · §Introduction
  • VirTex2020 · cited 2×
    “Other approaches use contrastive losses based on context prediction [20, 23], mutual information maximization [53, 54, 21], predicting masked regions [55], and clustering [56, 57, 58].”
    From VirTex · §Related Work
  • BYOL2020 · cited 2×
    “We follow the semi-supervised protocol of [74, 76, 8, 32] detailed in Section C.1, and use the same fixed splits of respectively 1%1\% and 10%10\% of ImageNet labeled training data as in [8].”
    From BYOL · §Experimental evaluation
  • ConVIRT2020 · cited 2×
    “Our work is inspired by the recent line of work on image view-based contrastive learning Hénaff et al. 2020; Chen et al. 2020a; He et al. 2020; Grill et al. 2020; Sowrirajan et al. 2021; Azizi et al. 2021, but fundamenta…”
    From ConVIRT · §Related Work
Abstract

Human observers can learn to recognize new categories of images from a handful of examples, yet doing so with artificial ones remains an open challenge. We hypothesize that data-efficient recognition is enabled by representations which make the variability in natural signals more predictable. We therefore revisit and improve Contrastive Predictive Coding, an unsupervised objective for learning such representations. This new implementation produces features which support state-of-the-art linear classification accuracy on the ImageNet dataset. When used as input for non-linear classification with deep neural networks, this representation allows us to use 2-5x less labels than classifiers trained directly on image pixels. Finally, this unsupervised representation substantially improves transfer learning to object detection on the PASCAL VOC dataset, surpassing fully supervised pre-trained ImageNet classifiers.