Paper Lineage
Esc
MethodJun 2020arXiv 2006.07733cs.LG

Bootstrap your own latent: A new approach to self-supervised Learning

Jean-Bastien Grill, Florian Strub, Florent Altché and 11 others

We introduce Bootstrap Your Own Latent (BYOL), a new approach to self-supervised image representation learning. BYOL relies on two neural networks, referred to as online and target networks, that interact and learn from each other.

From the abstract

Built on

10 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (10)

  • SimCLR2020 · cited 20×, 4 in Method
    “Most unsupervised methods for representation learning can be categorized as either generative or discriminative [23, 8].”
    From this paper · §Related work
  • MoCo2019 · cited 8×, 1 in Method
    “At the end of training, we only keep the encoder fθf_{\theta}; as in [9].”
    From this paper · §Method
  • CPC2018 · cited 5×, 1 in Method
    “We first evaluate BYOL’s representation by training a linear classifier on top of the frozen representation, following the procedure described in [48, 74, 41, 10, 8], and section C.1; we report top-11 and top-55 accuraci…”
    From this paper · §Experimental evaluation
  • ResNet2015 · cited 2×, 1 in Method
    “We use a convolutional residual network [22] with 50 layers and post-activation (ResNet-50(1×)50(1\times) v1) as our base parametric encoders fθf_{\theta} and fξf_{\xi}.”
    From this paper · §Method
Show 6 more
  • BatchNorm2015 · cited 1×, 1 in Method
    “This MLP consists in a linear layer with output size 40964096 followed by batch normalization [68], rectified linear units (ReLU) [69], and a final linear layer with output dimension 256256.”
    From this paper · §Method
  • Wide ResNet2016 · cited 1×, 1 in Method
    “We also use deeper (5050, 101101, 152152 and 200200 layers) and wider (from 1×1\times to 4×4\times) ResNets, as in [67, 48, 8].”
    From this paper · §Method
  • Goyal large-batch SGD2017 · cited 1×, 1 in Method
    “We set the base learning rate to 0.2,0.2, scaled linearly [72] with the batch size (LearningRate=0.2×BatchSize/256LearningRate=0.2\timesBatchSize/256).”
    From this paper · §Method
  • “We first evaluate BYOL’s representation by training a linear classifier on top of the frozen representation, following the procedure described in [48, 74, 41, 10, 8], and section C.1; we report top-11 and top-55 accuraci…”
    From this paper · §Experimental evaluation
  • VGG2014 · cited 2×
    “[1, 2, 3, 4, 5, 6, 7]”
    From this paper · §Introduction
  • CPC v22019 · cited 2×
    “We follow the semi-supervised protocol of [74, 76, 8, 32] detailed in Section C.1, and use the same fixed splits of respectively 1%1\% and 10%10\% of ImageNet labeled training data as in [8].”
    From this paper · §Experimental evaluation

Led to

  • ConVIRT2020 · cited 2×
    “Our work is inspired by the recent line of work on image view-based contrastive learning Hénaff et al. 2020; Chen et al. 2020a; He et al. 2020; Grill et al. 2020; Sowrirajan et al. 2021; Azizi et al. 2021, but fundamenta…”
    From ConVIRT · §Related Work
  • ViT-VQGAN2021 · cited 2×
    “In computer vision, in contrast, most recent unsupervised or self-supervised learning research focuses on applying different random augmentations to images, with the pretraining objective to distinguish image instances (…”
    From ViT-VQGAN · §Introduction
  • MAE2021 · cited 5×
    “Recently, contrastive learning Becker1992; Hadsell2006 has been popular, e.g., Wu2018a; Oord2018; He2020; Chen2020, which models image similarity and dissimilarity (or only similarity Grill2020; Chen2021) between two or…”
    From MAE · §Related Work
Abstract

We introduce Bootstrap Your Own Latent (BYOL), a new approach to self-supervised image representation learning. BYOL relies on two neural networks, referred to as online and target networks, that interact and learn from each other. From an augmented view of an image, we train the online network to predict the target network representation of the same image under a different augmented view. At the same time, we update the target network with a slow-moving average of the online network. While state-of-the art methods rely on negative pairs, BYOL achieves a new state of the art without them. BYOL reaches $74.3\%$ top-1 classification accuracy on ImageNet using a linear evaluation with a ResNet-50 architecture and $79.6\%$ with a larger ResNet. We show that BYOL performs on par or better than the current state of the art on both transfer and semi-supervised benchmarks. Our implementation and pretrained models are given on GitHub.