VirTex: Learning Visual Representations from Textual Annotations
The de-facto approach to many vision tasks is to start from pretrained visual representations, typically learned via supervised training on ImageNet. Recent methods have explored unsupervised pretraining to scale to vast quantities of unlabeled images.
Also cited · not yet reviewed (8)
- COCO Captions2015 · cited 4×, 1 in Method“Training Details: We train on the train2017 split of the COCO Captions dataset [36], which provides 118K118K images with five captions each.”From this paper · §Method
- RoBERTa2019 · cited 2×, 1 in Method“However, following BERT [64], many large-scale models [82, 83] instead use masked language models (MLMs): some tokens are randomly masked and are predicted by the model.”From this paper · §Method
- Megatron-LM2019 · cited 2×, 1 in Method“However, following BERT [64], many large-scale models [82, 83] instead use masked language models (MLMs): some tokens are randomly masked and are predicted by the model.”From this paper · §Method
- GPT-32020 · cited 2×, 1 in Method“This involves training massive language models – either unidirectional [77] or bidirectional [78, 79, 80, 81], for predicting tokens one by one.”From this paper · §Method
Show 4 more
- CPC2018 · cited 2דOther approaches use contrastive losses based on context prediction [20, 23], mutual information maximization [53, 54, 21], predicting masked regions [55], and clustering [56, 57, 58].”From this paper · §Related Work
- CPC v22019 · cited 2דOther approaches use contrastive losses based on context prediction [20, 23], mutual information maximization [53, 54, 21], predicting masked regions [55], and clustering [56, 57, 58].”From this paper · §Related Work
- VisualBERT2019 · cited 2דInspired by the success of BERT [64] in NLP, several recent methods use Transformers [29] to learn transferable joint representations of images and text [65, 66, 67, 68, 69, 70, 71, 72].”From this paper · §Related Work
- UNITER2019 · cited 2דInspired by the success of BERT [64] in NLP, several recent methods use Transformers [29] to learn transferable joint representations of images and text [65, 66, 67, 68, 69, 70, 71, 72].”From this paper · §Related Work
Led to
- ALIGN2021 · cited 2דJoulin et al. 2015; Li et al. 2017; Desai & Johnson 2020; Sariyildiz et al. 2020; Zhang et al. 2020 show that a good visual representation can be learned by predicting the captions from images, which inspires our work.”From ALIGN · §Related Work
- CLIP2021 · cited 2×, 1 in Method“Zhang et al. 2020, Gomez et al. 2017, Joulin et al. 2016, and Desai & Johnson 2020 all introduce methods which learn visual representations from text paired with images but describe their approaches as unsupervised, self…”From CLIP · §Approach
- Flamingo2022 · cited 1×, 1 in Method“In our ablation studies (Section 3.3), we compare the proposed gated xattn-dense layers against recent alternatives [22, 68] and explore the effect of how frequently these additional layers are inserted to trade off betw…”From Flamingo · §Approach
Abstract
The de-facto approach to many vision tasks is to start from pretrained visual representations, typically learned via supervised training on ImageNet. Recent methods have explored unsupervised pretraining to scale to vast quantities of unlabeled images. In contrast, we aim to learn high-quality visual representations from fewer images. To this end, we revisit supervised pretraining, and seek data-efficient alternatives to classification-based pretraining. We propose VirTex -- a pretraining approach using semantically dense captions to learn visual representations. We train convolutional networks from scratch on COCO Captions, and transfer them to downstream recognition tasks including image classification, object detection, and instance segmentation. On all tasks, VirTex yields features that match or exceed those learned on ImageNet -- supervised or unsupervised -- despite using up to ten times fewer images.