Paper Lineage
Esc

Transfer learning & weak supervision

Pre-train on huge, cheap or noisy labels, then transfer to downstream tasks.

Concepts in Transfer learning & weak supervision
ConceptIntroduced byYearsPapers
Off-the-shelf deep features
Activations of an ImageNet CNN make strong generic features for new tasks.
DeCAF2013–20141
Weakly supervised pre-training
Use huge amounts of cheap, noisy labels instead of clean annotations.
Weakly supervised visual features (Jouli2015–20213
Natural-language supervision for vision
Train image models to predict free-form words and phrases from accompanying text.
Visual N-Grams2016–20211
Data scaling in vision
Performance keeps rising (logarithmically) with 100× more labelled images.
JFT-300M (unreasonable effectiveness)2017–20215
Domain-adaptive pre-training data
Reweight pre-training data towards the target domain; more data is not always better.
Domain adaptive transfer20180
Hashtag supervision at billion scale
Pre-train on billions of social-media images labelled only by their hashtags.
Instagram hashtag pre-training2018–20226
Big Transfer recipe
Large-scale supervised pre-training plus a simple, fixed transfer heuristic.
BiT2019–20211
Noisy Student training
Self-training with an equal-or-larger student and noise added to it, iterated.
Noisy Student2019–20213
Teacher–student self-training
A teacher labels a huge unlabelled pool; a student learns from those pseudo-labels.
Billion-scale semi-supervised2019–20217
Captioning as visual pre-training
Learn visual backbones by generating captions, which is data-efficient compared to classification.
VirTex2020–20211