Paper Lineage
Esc
MethodJul 2018arXiv 1807.03748cs.LG

Representation Learning with Contrastive Predictive Coding

Aaron van den Oord, Yazhe Li, Oriol Vinyals

While supervised learning has enabled great progress in many applications, unsupervised learning has not seen such widespread adoption, and remains an important and challenging endeavor for artificial intelligence. In this work, we propose a universal unsupervised learning approach to extract useful representations from high-dimensional data, which we call Contrastive Predictive Coding.

From the abstract

Built on

1 paper · 0 verifiedSee as graph

Also cited · not yet reviewed (1)

  • word2vec2013 · cited 2×
    “Recent work in unsupervised learning has successfully used these ideas to learn word representations by predicting neighboring words [9].”
    From this paper · §Introduction

Led to

  • Deep InfoMax2018 · cited 6×
    “However, when we adopted the strided crop architecture found in Oord et al. 2018, both CPC and DIM performance improved considerably.”
    From Deep InfoMax · §Experiments
  • CPC v22019 · cited 8×, 3 in Method
    “This loss is called InfoNCE as it is inspired by Noise-Contrastive Estimation (Gutmann & Hyvärinen 2010; Mnih & Kavukcuoglu 2013) and has been shown to maximize the mutual information between 𝒄i,j{\bm{c}}_{i,j} and 𝒛i+…”
    From CPC v2 · §Experimental Setup
  • AMDIM2019 · cited 3×, 1 in Method
    “The best results with local DIM were obtained using a mutual information bound based on Noise-Contrastive Estimation (NCE – (Gutmann and Hyvärinen 2010)), as used in various NLP applications (Ma and Collins 2018), and ap…”
    From AMDIM · §Method Description
  • MoCo2019 · cited 12×, 7 in Method
    “With similarity measured by dot product, a form of a contrastive loss function, called InfoNCE Oord2018, is considered in this paper:”
    From MoCo · §Method
  • SimCLR2020 · cited 5×, 2 in Method
    “To evaluate the learned representations, we follow the widely used linear evaluation protocol (Zhang et al. 2016; Oord et al. 2018; Bachman et al. 2019; Kolesnikov et al. 2019), where a linear classifier is trained on to…”
    From SimCLR · §Method
  • VirTex2020 · cited 2×
    “Other approaches use contrastive losses based on context prediction [20, 23], mutual information maximization [53, 54, 21], predicting masked regions [55], and clustering [56, 57, 58].”
    From VirTex · §Related Work
  • BYOL2020 · cited 5×, 1 in Method
    “We first evaluate BYOL’s representation by training a linear classifier on top of the frozen representation, following the procedure described in [48, 74, 41, 10, 8], and section C.1; we report top-11 and top-55 accuraci…”
    From BYOL · §Experimental evaluation
  • ConVIRT2020 · cited 1×, 1 in Method
    “This loss takes the same form as the InfoNCE loss Oord et al. 2018, and minimizing it leads to encoders that maximally preserve the mutual information between the true pairs under the representation functions.”
    From ConVIRT · §Methods
  • CLIP2021 · cited 1×, 1 in Method
    “To our knowledge this batch construction technique and objective was first introduced in the area of deep metric learning as the multi-class N-pair loss Sohn 2016, was popularized for contrastive representation learning…”
    From CLIP · §Approach
Abstract

While supervised learning has enabled great progress in many applications, unsupervised learning has not seen such widespread adoption, and remains an important and challenging endeavor for artificial intelligence. In this work, we propose a universal unsupervised learning approach to extract useful representations from high-dimensional data, which we call Contrastive Predictive Coding. The key insight of our model is to learn such representations by predicting the future in latent space by using powerful autoregressive models. We use a probabilistic contrastive loss which induces the latent space to capture information that is maximally useful to predict future samples. It also makes the model tractable by using negative sampling. While most prior work has focused on evaluating representations for a particular modality, we demonstrate that our approach is able to learn useful representations achieving strong performance on four distinct domains: speech, images, text and reinforcement learning in 3D environments.