Paper Lineage
Esc
MethodNov 2019arXiv 1911.05722cs.CV

Momentum Contrast for Unsupervised Visual Representation Learning

Kaiming He, Haoqi Fan, Yuxin Wu and 2 others

We present Momentum Contrast (MoCo) for unsupervised visual representation learning. From a perspective on contrastive learning as dictionary look-up, we build a dynamic dictionary with a queue and a moving-averaged encoder.

From the abstract

Built on

12 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (12)

  • Instance discrimination2018 · cited 23×, 11 in Method
    “In this paper, we follow a simple instance discrimination task Wu2018a; Ye2019; Bachman2019: a query matches a key if they are encoded views (e.g., different crops) of the same image.”
    From this paper · §Introduction
  • CPC2018 · cited 12×, 7 in Method
    “With similarity measured by dot product, a form of a contrastive loss function, called InfoNCE Oord2018, is considered in this paper:”
    From this paper · §Method
  • AMDIM2019 · cited 10×, 6 in Method
    “In this paper, we follow a simple instance discrimination task Wu2018a; Ye2019; Bachman2019: a query matches a key if they are encoded views (e.g., different crops) of the same image.”
    From this paper · §Introduction
  • Invariant & spreading instance features2019 · cited 6×, 5 in Method
    “As the focus of this paper is not on designing a new pretext task, we use a simple one mainly following the instance discrimination task in Wu2018a, to which some recent works Ye2019; Bachman2019 are related.”
    From this paper · §Method
Show 8 more
  • Deep InfoMax2018 · cited 6×, 4 in Method
    “Several recent studies Wu2018a; Oord2018; Hjelm2019; Zhuang2019; Henaff2019; Tian2019; Bachman2019 present promising results on unsupervised visual representation learning using approaches related to the contrastive loss…”
    From this paper · §Introduction
  • CPC v22019 · cited 5×, 2 in Method
    “Several recent studies Wu2018a; Oord2018; Hjelm2019; Zhuang2019; Henaff2019; Tian2019; Bachman2019 present promising results on unsupervised visual representation learning using approaches related to the contrastive loss…”
    From this paper · §Introduction
  • ResNet2015 · cited 4×, 2 in Method
    “We adopt a ResNet He2016 as the encoder, whose last fully-connected layer (after global average pooling) has a fixed-dimensional output (128-D Wu2018a).”
    From this paper · §Method
  • Goyal large-batch SGD2017 · cited 4×, 1 in Method
    “It is also challenged by large mini-batch optimization Goyal2017.”
    From this paper · §Method
  • BatchNorm2015 · cited 1×, 1 in Method
    “Our encoders fqf_{\textrm{q}} and fkf_{\textrm{k}} both have Batch Normalization (BN) Ioffe2015 as in the standard ResNet He2016.”
    From this paper · §Method
  • Faster R-CNN2015 · cited 2×
    “The detector is Faster R-CNN Ren2015 with a backbone of R50-dilated-C5 or R50-C4 He2017 (details in appendix), with BN tuned, implemented in Wu2019.”
    From this paper · §Experiments
  • “Instagram-1B (IG-1B): Following Mahajan2018, this is a dataset of ∼\scriptstyle\sim1 billion (940M) public images from Instagram.”
    From this paper · §Experiments
  • BERT2018 · cited 2×
    “Beyond the simple instance discrimination task Wu2018a, it is possible to adopt MoCo for pretext tasks like masked auto-encoding, e.g., in language Devlin2019 and in vision Oord2018.”
    From this paper · §Discussion and Conclusion

Led to

  • SimCLR2020 · cited 6×, 2 in Method
    “To keep it simple, we do not train the model with a memory bank (Wu et al. 2018; He et al. 2019).”
    From SimCLR · §Method
  • BYOL2020 · cited 8×, 1 in Method
    “At the end of training, we only keep the encoder fθf_{\theta}; as in [9].”
    From BYOL · §Method
  • ConVIRT2020 · cited 5×, 1 in Method
    “Note that unlike previous work which use a contrastive loss between inputs of the same modality Chen et al. 2020a; He et al. 2020, our image-to-text contrastive loss is asymmetric for each input modality.”
    From ConVIRT · §Methods
  • BEiT2021 · cited 2×
    “The recent strand of research follows contrastive paradigm [43, 31, 16, 3, 17, 7, 5].”
    From BEiT · §Related Work
  • ALBEF2021 · cited 3×, 1 in Method
    “Inspired by MoCo [24], we maintain two queues to store the most recent MM image-text representations from the momentum unimodal encoders.”
    From ALBEF · §ALBEF Pre-training
  • MAE2021 · cited 2×
    “Recently, contrastive learning Becker1992; Hadsell2006 has been popular, e.g., Wu2018a; Oord2018; He2020; Chen2020, which models image similarity and dissimilarity (or only similarity Grill2020; Chen2021) between two or…”
    From MAE · §Related Work
Abstract

We present Momentum Contrast (MoCo) for unsupervised visual representation learning. From a perspective on contrastive learning as dictionary look-up, we build a dynamic dictionary with a queue and a moving-averaged encoder. This enables building a large and consistent dictionary on-the-fly that facilitates contrastive unsupervised learning. MoCo provides competitive results under the common linear protocol on ImageNet classification. More importantly, the representations learned by MoCo transfer well to downstream tasks. MoCo can outperform its supervised pre-training counterpart in 7 detection/segmentation tasks on PASCAL VOC, COCO, and other datasets, sometimes surpassing it by large margins. This suggests that the gap between unsupervised and supervised representation learning has been largely closed in many vision tasks.