Paper Lineage
Esc

Self-supervised visual learning

Learning visual features without labels: contrastive, mutual-information and masked-prediction methods.

Concepts in Self-supervised visual learning
ConceptIntroduced byYearsPapers
Contrastive predictive coding / InfoNCE
Predict future latents and score them against negatives with a contrastive loss.
CPC2018–20216
Instance discrimination
Treat every image as its own class and learn features that tell instances apart.
Instance discrimination2018–20191
Mutual-information maximisation
Learn features by maximising information shared between an input and its encoding, or between views.
Deep InfoMax2018–20214
Momentum contrast
A queue of negatives encoded by a slowly-updated momentum encoder.
MoCo2019–20212
Augmentations + projection head
Strong augmentation plus a nonlinear projection head is what makes contrastive learning work.
SimCLR20201
Self-distillation without negatives
An online network predicts a slowly-moving target network's view, with no negatives needed.
BYOL2020–20211
Masked autoencoders
Mask 75% of patches, encode only the visible ones, and reconstruct pixels with a light decoder.
MAE2021–20222
Masked image modelling
Hide image patches and predict them, the vision analogue of BERT.
BEiT2021–20224
Masked prediction of CLIP features
Reconstruct masked-out CLIP features, which scales a plain ViT to a billion parameters.
EVA2022–20231
Semantic visual tokenizer (VQ-KD)
Distil a semantic teacher into discrete codes, and use them as masked-prediction targets.
BEiT v220221