Transfer learning & weak supervision
Pre-train on huge, cheap or noisy labels, then transfer to downstream tasks.
| Concept | Introduced by | Years | Papers |
|---|---|---|---|
| Off-the-shelf deep features Activations of an ImageNet CNN make strong generic features for new tasks. | DeCAF | 2013–2014 | 1 |
| Weakly supervised pre-training Use huge amounts of cheap, noisy labels instead of clean annotations. | Weakly supervised visual features (Jouli | 2015–2021 | 3 |
| Natural-language supervision for vision Train image models to predict free-form words and phrases from accompanying text. | Visual N-Grams | 2016–2021 | 1 |
| Data scaling in vision Performance keeps rising (logarithmically) with 100× more labelled images. | JFT-300M (unreasonable effectiveness) | 2017–2021 | 5 |
| Domain-adaptive pre-training data Reweight pre-training data towards the target domain; more data is not always better. | Domain adaptive transfer | 2018 | 0 |
| Hashtag supervision at billion scale Pre-train on billions of social-media images labelled only by their hashtags. | Instagram hashtag pre-training | 2018–2022 | 6 |
| Big Transfer recipe Large-scale supervised pre-training plus a simple, fixed transfer heuristic. | BiT | 2019–2021 | 1 |
| Noisy Student training Self-training with an equal-or-larger student and noise added to it, iterated. | Noisy Student | 2019–2021 | 3 |
| Teacher–student self-training A teacher labels a huge unlabelled pool; a student learns from those pseudo-labels. | Billion-scale semi-supervised | 2019–2021 | 7 |
| Captioning as visual pre-training Learn visual backbones by generating captions, which is data-efficient compared to classification. | VirTex | 2020–2021 | 1 |