A Simple Framework for Contrastive Learning of Visual Representations
This paper presents SimCLR: a simple framework for contrastive learning of visual representations. We simplify recently proposed contrastive self-supervised learning algorithms without requiring specialized architectures or a memory bank.
Also cited · not yet reviewed (11)
- Instance discrimination2018 · cited 6×, 3 in Method“To keep it simple, we do not train the model with a memory bank (Wu et al. 2018; He et al. 2019).”From this paper · §Method
- AMDIM2019 · cited 9×, 2 in Method“To evaluate the learned representations, we follow the widely used linear evaluation protocol (Zhang et al. 2016; Oord et al. 2018; Bachman et al. 2019; Kolesnikov et al. 2019), where a linear classifier is trained on to…”From this paper · §Method
- MoCo2019 · cited 6×, 2 in Method“To keep it simple, we do not train the model with a memory bank (Wu et al. 2018; He et al. 2019).”From this paper · §Method
- CPC2018 · cited 5×, 2 in Method“To evaluate the learned representations, we follow the widely used linear evaluation protocol (Zhang et al. 2016; Oord et al. 2018; Bachman et al. 2019; Kolesnikov et al. 2019), where a linear classifier is trained on to…”From this paper · §Method
Show 7 more
- ResNet2015 · cited 2×, 2 in Method“We opt for simplicity and adopt the commonly used ResNet (He et al. 2016) to obtain 𝒉i=f(𝒙~i)=ResNet(𝒙~i)\bm{h}_{i}=f(\tilde{\bm{x}}_{i})=ResNet(\tilde{\bm{x}}_{i}) where 𝒉i∈ℝd\bm{h}_{i}\in\mathbb{R}^{d} is the out…”From this paper · §Method
- CPC v22019 · cited 10×, 1 in Method“Not only does SimCLR outperform previous work (Figure 1), but it is also simpler, requiring neither specialized architectures (Bachman et al. 2019; Hénaff et al. 2019) nor a memory bank (Wu et al. 2018; Tian et al. 2019;…”From this paper · §Introduction
- Goyal large-batch SGD2017 · cited 2×, 1 in Method“In contrast to supervised learning (Goyal et al. 2017), in contrastive learning, larger batch sizes provide more negative examples, facilitating convergence (i.e. taking fewer epochs and steps for a given accuracy).”From this paper · §Loss Functions and Batch Size
- BatchNorm2015 · cited 1×, 1 in Method“Standard ResNets use batch normalization (Ioffe & Szegedy 2015).”From this paper · §Method
- GoogLeNet (Inception)2014 · cited 2דThe other type of augmentation involves appearance transformation, such as color distortion (including color dropping, brightness, contrast, saturation, hue) (Howard 2013; Szegedy et al. 2015), Gaussian blur, and Sobel f…”From this paper · §Data Augmentation for Contrastive Representation Learning
- Do better ImageNet models transfer bette2018 · cited 2דFollowing Kornblith et al. 2019, we perform hyperparameter tuning for each model-dataset combination and select the best hyperparameters on a validation set.”From this paper · §Comparison with State-of-the-art
- Deep InfoMax2018 · cited 2דRecent literature has attempted to relate the success of their methods to maximization of mutual information between latent representations (Oord et al. 2018; Hénaff et al. 2019; Hjelm et al. 2018; Bachman et al. 2019).”From this paper · §Related Work
Led to
- BYOL2020 · cited 20×, 4 in Method“Most unsupervised methods for representation learning can be categorized as either generative or discriminative [23, 8].”From BYOL · §Related work
- ConVIRT2020 · cited 9×, 3 in Method“Note that unlike previous work which use a contrastive loss between inputs of the same modality Chen et al. 2020a; He et al. 2020, our image-to-text contrastive loss is asymmetric for each input modality.”From ConVIRT · §Methods
- ViLT2021 · cited 2דPrevious work on contrastive visual representation learning (Chen et al. 2020a; Chen et al. 2020b) showed that gaussian blur, not employed by RandAugment, brings noticeable gains to downstream performance compared with a…”From ViLT · §Conclusion and Future Work
- ALIGN2021 · cited 2דRecently, self-supervised (Chen et al. 2020b; Tian et al. 2020; He et al. 2020; Misra & Maaten 2020; Li et al. 2021; Grill et al. 2020; Caron et al. 2020) and semi-supervised learning (Yalniz et al. 2019; Xie et al. 2020…”From ALIGN · §Related Work
- CLIP2021 · cited 1×, 1 in Method“We do not use the non-linear projection between the representation and the contrastive embedding space, a change which was introduced by Bachman et al. 2019 and popularized by Chen et al. 2020b.”From CLIP · §Approach
- BEiT2021 · cited 2דThe recent strand of research follows contrastive paradigm [43, 31, 16, 3, 17, 7, 5].”From BEiT · §Related Work
- ALBEF2021 · cited 2דThe recent CLIP [6] and ALIGN [7] perform pre-training on massive noisy web data using a contrastive loss, one of the most effective loss for representation learning [24, 25, 26, 27].”From ALBEF · §Related Work
- MAE2021 · cited 4דRecently, contrastive learning Becker1992; Hadsell2006 has been popular, e.g., Wu2018a; Oord2018; He2020; Chen2020, which models image similarity and dissimilarity (or only similarity Grill2020; Chen2021) between two or…”From MAE · §Related Work
- UniCL2022 · cited 2דContrastive learning has laid the foundation for the best performing SSL models tian2019contrastive; henaff2019data; chen2020simple; he2020momentum; caron2020unsupervised; tian2020makes; chen2021empirical.”From UniCL · §Related works
Abstract
This paper presents SimCLR: a simple framework for contrastive learning of visual representations. We simplify recently proposed contrastive self-supervised learning algorithms without requiring specialized architectures or a memory bank. In order to understand what enables the contrastive prediction tasks to learn useful representations, we systematically study the major components of our framework. We show that (1) composition of data augmentations plays a critical role in defining effective predictive tasks, (2) introducing a learnable nonlinear transformation between the representation and the contrastive loss substantially improves the quality of the learned representations, and (3) contrastive learning benefits from larger batch sizes and more training steps compared to supervised learning. By combining these findings, we are able to considerably outperform previous methods for self-supervised and semi-supervised learning on ImageNet. A linear classifier trained on self-supervised representations learned by SimCLR achieves 76.5% top-1 accuracy, which is a 7% relative improvement over previous state-of-the-art, matching the performance of a supervised ResNet-50. When fine-tuned on only 1% of the labels, we achieve 85.8% top-5 accuracy, outperforming AlexNet with 100X fewer labels.