Self-supervised visual learning
Learning visual features without labels: contrastive, mutual-information and masked-prediction methods.
| Concept | Introduced by | Years | Papers |
|---|---|---|---|
| Contrastive predictive coding / InfoNCE Predict future latents and score them against negatives with a contrastive loss. | CPC | 2018–2021 | 6 |
| Instance discrimination Treat every image as its own class and learn features that tell instances apart. | Instance discrimination | 2018–2019 | 1 |
| Mutual-information maximisation Learn features by maximising information shared between an input and its encoding, or between views. | Deep InfoMax | 2018–2021 | 4 |
| Momentum contrast A queue of negatives encoded by a slowly-updated momentum encoder. | MoCo | 2019–2021 | 2 |
| Augmentations + projection head Strong augmentation plus a nonlinear projection head is what makes contrastive learning work. | SimCLR | 2020 | 1 |
| Self-distillation without negatives An online network predicts a slowly-moving target network's view, with no negatives needed. | BYOL | 2020–2021 | 1 |
| Masked autoencoders Mask 75% of patches, encode only the visible ones, and reconstruct pixels with a light decoder. | MAE | 2021–2022 | 2 |
| Masked image modelling Hide image patches and predict them, the vision analogue of BERT. | BEiT | 2021–2022 | 4 |
| Masked prediction of CLIP features Reconstruct masked-out CLIP features, which scales a plain ViT to a billion parameters. | EVA | 2022–2023 | 1 |
| Semantic visual tokenizer (VQ-KD) Distil a semantic teacher into discrete codes, and use them as masked-prediction targets. | BEiT v2 | 2022 | 1 |