Learning Transferable Visual Models From Natural Language Supervision
State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories. This restricted form of supervision limits their generality and usability since additional labeled data is needed to specify any other visual concept.
Also cited · not yet reviewed (27)
- ConVIRT2020 · cited 5×, 4 in Method“Zhang et al. 2020, Gomez et al. 2017, Joulin et al. 2016, and Desai & Johnson 2020 all introduce methods which learn visual representations from text paired with images but describe their approaches as unsupervised, self…”From this paper · §Approach
- Instagram hashtag pre-training2018 · cited 9×, 3 in Method“By comparison, other computer vision systems are trained on up to 3.5 billion Instagram photos (Mahajan et al. 2018).”From this paper · §Approach
- ResNet2015 · cited 2×, 2 in Method“For the first, we use ResNet-50 (He et al. 2016a) as the base architecture for the image encoder due to its widespread adoption and proven performance.”From this paper · §Approach
- EfficientNet2019 · cited 2×, 2 in Method“While previous computer vision research has often scaled models by increasing the width (Mahajan et al. 2018) or depth (He et al. 2016a) in isolation, for the ResNet image encoders we adapt the approach of Tan & Le 2019…”From this paper · §Approach
Show 23 more
- ViT2020 · cited 4×, 1 in Method“For the second architecture, we experiment with the recently introduced Vision Transformer (ViT) (Dosovitskiy et al. 2020).”From this paper · §Approach
- Noisy Student2019 · cited 3×, 1 in Method“Mahajan et al. 2018 required 19 GPU years to train their ResNeXt101-32x48d and Xie et al. 2020 required 33 TPUv3 core-years to train their Noisy Student EfficientNet-L2.”From this paper · §Approach
- YFCC100M2015 · cited 2×, 1 in Method“Existing work has mainly used three datasets, MS-COCO (Lin et al. 2014), Visual Genome (Krishna et al. 2017), and YFCC100M (Thomee et al. 2016).”From this paper · §Approach
- Weakly supervised visual features (Jouli2015 · cited 2×, 1 in Method“Zhang et al. 2020, Gomez et al. 2017, Joulin et al. 2016, and Desai & Johnson 2020 all introduce methods which learn visual representations from text paired with images but describe their approaches as unsupervised, self…”From this paper · §Approach
- VirTex2020 · cited 2×, 1 in Method“Zhang et al. 2020, Gomez et al. 2017, Joulin et al. 2016, and Desai & Johnson 2020 all introduce methods which learn visual representations from text paired with images but describe their approaches as unsupervised, self…”From this paper · §Approach
- MS COCO2014 · cited 1×, 1 in Method“Existing work has mainly used three datasets, MS-COCO (Lin et al. 2014), Visual Genome (Krishna et al. 2017), and YFCC100M (Thomee et al. 2016).”From this paper · §Approach
- Visual Genome2016 · cited 1×, 1 in Method“Existing work has mainly used three datasets, MS-COCO (Lin et al. 2014), Visual Genome (Krishna et al. 2017), and YFCC100M (Thomee et al. 2016).”From this paper · §Approach
- Transformer2017 · cited 1×, 1 in Method“The text encoder is a Transformer (Vaswani et al. 2017) with the architecture modifications described in Radford et al. 2019.”From this paper · §Approach
- Learned in Translation2017 · cited 1×, 1 in Method“Although early work wrestled with the complexity of natural language when using topic model and n-gram representations, improvements in deep contextual representation learning suggest we now have the tools to effectively…”From this paper · §Approach
- AdamW2017 · cited 1×, 1 in Method“We use the Adam optimizer (Kingma & Ba 2014) with decoupled weight decay regularization (Loshchilov & Hutter 2017) applied to all weights that are not gains or biases, and decay the learning rate using a cosine schedule…”From this paper · §Approach
- Instance discrimination2018 · cited 1×, 1 in Method“The learnable temperature parameter τ\tau was initialized to the equivalent of 0.07 from (Wu et al. 2018) and clipped to prevent scaling the logits by more than 100 which we found necessary to prevent training instabilit…”From this paper · §Approach
- CPC2018 · cited 1×, 1 in Method“To our knowledge this batch construction technique and objective was first introduced in the area of deep metric learning as the multi-class N-pair loss Sohn 2016, was popularized for contrastive representation learning…”From this paper · §Approach
- AMDIM2019 · cited 1×, 1 in Method“We do not use the non-linear projection between the representation and the contrastive embedding space, a change which was introduced by Bachman et al. 2019 and popularized by Chen et al. 2020b.”From this paper · §Approach
- SimCLR2020 · cited 1×, 1 in Method“We do not use the non-linear projection between the representation and the contrastive embedding space, a change which was introduced by Bachman et al. 2019 and popularized by Chen et al. 2020b.”From this paper · §Approach
- Natural distribution shift robustness2020 · cited 8דEncouragingly however, Taori et al. 2020 find that accuracy under distribution shift increases predictably with ImageNet accuracy and is well modeled as a linear function of logit-transformed accuracy.”From this paper · §Experiments
- Do better ImageNet models transfer bette2018 · cited 6דThis increases flexibility, and prior work has convincingly demonstrated that fine-tuning outperforms linear classification on most image classification datasets (Kornblith et al. 2019; Zhai et al. 2019).”From this paper · §Experiments
- BiT2019 · cited 5דKolesnikov et al. 2019 and Dosovitskiy et al. 2020 have also demonstrated large gains on a broader set of transfer benchmarks by pre-training models to predict the classes of the noisily labeled JFT-300M dataset.”From this paper · §Introduction and Motivating Work
- Visual N-Grams2016 · cited 4דLi et al. 2017 then extended this approach to predicting phrase n-grams in addition to individual words and demonstrated the ability of their system to zero-shot transfer to other image classification datasets by scoring…”From this paper · §Introduction and Motivating Work
- GPT-32020 · cited 4דSimilar to the “prompt engineering” discussion around GPT-3 (Brown et al. 2020; Gao et al. 2020), we have also observed that zero-shot performance can be significantly improved by customizing the prompt text to each task…”From this paper · §Experiments
- ImageNet-trained CNNs are biased towards2018 · cited 2דHowever, research in the subsequent years has repeatedly found that these models still make many simple mistakes (Dodge & Karam 2017; Geirhos et al. 2018; Alcorn et al. 2019), and new benchmarks testing these systems has…”From this paper · §Experiments
- ImageNetV22019 · cited 2דHowever, research in the subsequent years has repeatedly found that these models still make many simple mistakes (Dodge & Karam 2017; Geirhos et al. 2018; Alcorn et al. 2019), and new benchmarks testing these systems has…”From this paper · §Experiments
- T52019 · cited 2דPre-training methods which learn directly from raw text have revolutionized NLP over the last few years (Dai & Le 2015; Peters et al. 2018; Howard & Ruder 2018; Radford et al. 2018; Devlin et al. 2018; Raffel et al. 2019…”From this paper · §Introduction and Motivating Work
- Kaplan scaling laws2020 · cited 2דWe study the scalability of CLIP by training a series of eight models spanning almost 2 orders of magnitude of compute and observe that transfer performance is a smoothly predictable function of compute (Hestness et al.…”From this paper · §Introduction and Motivating Work
Led to
- VL-T52021 · cited 2דFollowing this success, image+text pretraining models (Lu et al. 2019; Tan & Bansal 2019; Chen et al. 2020; Huang et al. 2020; Li et al. 2020b; Cho et al. 2020; Radford et al. 2021; Zhang et al. 2021) and video+text pret…”From VL-T5 · §Related Works
- ViLT2021 · cited 1×, 1 in Method“CLIP (Radford et al. 2021) belongs to Figure 2(b) as it uses separate but equally expensive transformer embedders for each modality.”From ViLT · §Background
- ALIGN2021 · cited 2דIn the zero-shot setting, ALIGN gets more than 7% improvement in image retrieval task compared to the previous SOTA, CLIP (Radford et al. 2021).”From ALIGN · §Experiments and Results
- DALL·E2021 · cited 2×, 1 in Method“Similar to Razavi et al. 2019, we rerank the samples drawn from the transformer using a pretrained contrastive model (Radford et al. 2021).”From DALL·E · §Method
- True few-shot learning2021 · cited 2דPrior work uses large train or held-out sets with many examples to choose prompts [2, 12, 13] and hyperparameters [12].”From True few-shot learning · §Introduction
- ALBEF2021 · cited 4דThe first category focuses on modelling the interactions between image and text features with transformer-based multimodal encoders [10, 11, 12, 13, 1, 14, 15, 2, 3, 16, 8, 17, 18].”From ALBEF · §Related Work
- SimVLM2021 · cited 4דMotivated by recent works (Radford et al. 2021; Ramesh et al. 2021; Jia et al. 2021; Tsimpoukelli et al. 2021) that illustrate zero-shot learning in certain image-text tasks, we train our model using large-scale weakly l…”From SimVLM · §Related Work
- LAION-400M2021 · cited 3דMulti-modal language-vision models demonstrated recently strong transfer capability to novel datasets in absense of per-sample labels [1, 2, 3].”From LAION-400M · §Introduction
- VLMo2021 · cited 4דPre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”From VLMo · §Related Work
- METER2021 · cited 3×, 1 in Method“Specifically, as shown in Figure 1, we dissect the model designs along multiple dimensions, including vision encoders (e.g., CLIP-ViT radford2021learning, Swin transformer liu2021swin), text encoders (e.g., RoBERTa liu20…”From METER · §Introduction
- LiT2021 · cited 17×, 4 in Method“Recently it was demonstrated that web-sourced paired image-text data can be used to pre-train strong models for zero-shot transfer clip; align.”From LiT · §Introduction
- Florence2021 · cited 16×, 3 in Method“In addition, we follow the sampling strategy introduced in (Radford et al. 2021; Ramesh et al. 2021) with the goal of achieving improved balance, informativeness, and learnability of the sampled dataset.”From Florence · §Approach
- BLIP2022 · cited 5×, 1 in Method“Due to the prohibitive expense of acquiring human-annotated texts, most methods (Chen et al. 2020; Li et al. 2020; Li et al. 2021a; Wang et al. 2021; Radford et al. 2021) use image and alt-text pairs crawled from the web…”From BLIP · §Related Work
- Simple end-to-end captioning2022 · cited 3×, 1 in Method“Visual Encoder.We follow the same setting as the visual transformer (ViT) which recently attracts much attention and shows remarkable performance in the computer vision area (Dosovitskiy et al. 2021; Radford et al. 2021;…”From Simple end-to-end captioning · §. Methodology
- UniCL2022 · cited 14×, 1 in Method“Typically, this can be tackled via either supervised learning on human-annotated image-label pairs deng2009imagenet or contrastive learning on webly-crawed image-text pairs radford2021learning; jia2021scaling.”From UniCL · §Introduction
- Flamingo2022 · cited 5×, 1 in Method“We pretrain the vision encoder using a contrastive objective on our datasets of image and text pairs, using the two-term contrastive loss from Radford et al. 2021.”From Flamingo · §Approach
- CoCa2022 · cited 13×, 2 in Method“Following previous practices [12, 32], “zero-shot” here is different from classical zero-shot learning in that during pretraining, the model may see relevant supervised information, but no supervised examples are used du…”From CoCa · §Approach
- VL-BEiT2022 · cited 4דDual-encoder model [29, 14] consists of an image encoder and a text encoder.”From VL-BEiT · §Related Work
- BEiT v22022 · cited 3×, 1 in Method“The output vectors {𝒐i}i=1N\{{\bm{o}}_{i}\}_{i=1}^{N} aim at reconstructing the semantic features of a teacher model, e.g., DINO (Caron et al. 2021), and CLIP (Radford et al. 2021).”From BEiT v2 · §Methodology
- BEiT-32022 · cited 5×, 3 in Method“In comparison, contrastive-based models [43, 23, 60, 62] usually need a very large batch size11 1 For example, CoCa [62] uses 6565k batch size, CLIP [43] uses 3232k batch size, and Florence [60] uses 2424k batch size.”From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
- EVA2022 · cited 5דRecently, there are a few trials leveraging the semantic information from image-image or image-text contrastive learning caron2021emerging; chen2021mocov3; clip for MIM pre-training zhou2021ibot; wei2022mvp; hou2022milan…”From EVA · §Introduction
- BLIP-22023 · cited 5×, 1 in Method“For the frozen image encoder, we explore two state-of-the-art pre-trained vision transformer models: (1) ViT-L/14 from CLIP (Radford et al. 2021) and (2) ViT-g/14 from EVA-CLIP (Fang et al. 2022).”From BLIP-2 · §Method
Abstract
State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories. This restricted form of supervision limits their generality and usability since additional labeled data is needed to specify any other visual concept. Learning directly from raw text about images is a promising alternative which leverages a much broader source of supervision. We demonstrate that the simple pre-training task of predicting which caption goes with which image is an efficient and scalable way to learn SOTA image representations from scratch on a dataset of 400 million (image, text) pairs collected from the internet. After pre-training, natural language is used to reference learned visual concepts (or describe new ones) enabling zero-shot transfer of the model to downstream tasks. We study the performance of this approach by benchmarking on over 30 different existing computer vision datasets, spanning tasks such as OCR, action recognition in videos, geo-localization, and many types of fine-grained object classification. The model transfers non-trivially to most tasks and is often competitive with a fully supervised baseline without the need for any dataset specific training. For instance, we match the accuracy of the original ResNet-50 on ImageNet zero-shot without needing to use any of the 1.28 million training examples it was trained on. We release our code and pre-trained model weights at https://github.com/OpenAI/CLIP.