Paper Lineage
Esc
MethodJun 2019arXiv 1906.00910cs.LG

Learning Representations by Maximizing Mutual Information Across Views

Philip Bachman, R Devon Hjelm, William Buchwalter

We propose an approach to self-supervised representation learning based on maximizing mutual information between features extracted from multiple views of a shared context. , tactile, auditory, or visual).

From the abstract

Built on

4 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (4)

  • CPC2018 · cited 3×, 1 in Method
    “The best results with local DIM were obtained using a mutual information bound based on Noise-Contrastive Estimation (NCE – (Gutmann and Hyvärinen 2010)), as used in various NLP applications (Ma and Collins 2018), and ap…”
    From this paper · §Method Description
  • Deep InfoMax2018 · cited 2×, 1 in Method
    “We introduce a model for self-supervised representation learning based on local Deep InfoMax (Hjelm et al. 2019, DIM,).”
    From this paper · §Introduction
  • ResNet2015 · cited 1×, 1 in Method
    “Our model uses an encoder based on the standard ResNet (He et al. 2016a; He et al. 2016b), with changes to make it suitable for DIM.”
    From this paper · §Method Description
  • BERT2018 · cited 2×
    “Self-supervised learning is gaining popularity across the NLP, vision, and robotics communities – e.g., (Devlin et al. 2019; Logeswaran and Lee 2018; Sermanet et al. 2017; Dwibedi et al. 2018).”
    From this paper · §Related Work

Led to

  • CPC v22019 · cited 2×
    “Augmented Multiscale Deep InfoMax (AMDIM, Bachman et al. 2019) is most similar to CPC in that it makes predictions across space, but differs in that it also predicts representations across layers in the model.”
    From CPC v2 · §Related Work
  • MoCo2019 · cited 10×, 6 in Method
    “In this paper, we follow a simple instance discrimination task Wu2018a; Ye2019; Bachman2019: a query matches a key if they are encoded views (e.g., different crops) of the same image.”
    From MoCo · §Introduction
  • SimCLR2020 · cited 9×, 2 in Method
    “To evaluate the learned representations, we follow the widely used linear evaluation protocol (Zhang et al. 2016; Oord et al. 2018; Bachman et al. 2019; Kolesnikov et al. 2019), where a linear classifier is trained on to…”
    From SimCLR · §Method
  • CLIP2021 · cited 1×, 1 in Method
    “We do not use the non-linear projection between the representation and the contrastive embedding space, a change which was introduced by Bachman et al. 2019 and popularized by Chen et al. 2020b.”
    From CLIP · §Approach
Abstract

We propose an approach to self-supervised representation learning based on maximizing mutual information between features extracted from multiple views of a shared context. For example, one could produce multiple views of a local spatio-temporal context by observing it from different locations (e.g., camera positions within a scene), and via different modalities (e.g., tactile, auditory, or visual). Or, an ImageNet image could provide a context from which one produces multiple views by repeatedly applying data augmentation. Maximizing mutual information between features extracted from these views requires capturing information about high-level factors whose influence spans multiple views -- e.g., presence of certain objects or occurrence of certain events. Following our proposed approach, we develop a model which learns image representations that significantly outperform prior methods on the tasks we consider. Most notably, using self-supervised learning, our model learns representations which achieve 68.1% accuracy on ImageNet using standard linear evaluation. This beats prior results by over 12% and concurrent results by 7%. When we extend our model to use mixture-based representations, segmentation behaviour emerges as a natural side-effect. Our code is available online: https://github.com/Philip-Bachman/amdim-public.