LiT: Zero-Shot Transfer with Locked-image text Tuning
This paper presents contrastive-tuning, a simple method employing contrastive training to align image and text models while still taking advantage of their pre-training. In our empirical study we find that locked pre-trained image models with unlocked text models work best.
Also cited · not yet reviewed (12)
- CLIP2021 · cited 17×, 4 in Method“Recently it was demonstrated that web-sourced paired image-text data can be used to pre-train strong models for zero-shot transfer clip; align.”From this paper · §Introduction
- ALIGN2021 · cited 14×, 3 in Method“Recently it was demonstrated that web-sourced paired image-text data can be used to pre-train strong models for zero-shot transfer clip; align.”From this paper · §Introduction
- ConVIRT2020 · cited 2×, 1 in Method“Another approach, which we adopt in this work, is to learn an alignment between image and text embedding spaces devise; vse; karpathy; vse-pooling; virtex; convirt.”From this paper · §Related work
- JFT-300M (unreasonable effectiveness)2017 · cited 1×, 1 in Method“Some dataset choices for learning powerful image embeddings are ImageNet-21k imagenet, JFT-300M unreasonable_effectiveness_of_data.”From this paper · §Methods
Show 8 more
- Scaling ViTs (ViT-G)2021 · cited 8דWith the pre-trained model ViT-g/14 vitg, LiT achieves 85.2% zero-shot transfer accuracy on ImageNet, halving the gap between previous best zero-shot transfer results clip; align and supervised fine-tuning results vitg;…”From this paper · §Introduction
- BiT2019 · cited 6דTransfer learning transfer_learning_survey has been a successful paradigm in computer vision imagenet_transfer_better; bit; instagram_resnext.”From this paper · §Introduction
- ViT2020 · cited 5דWe verify LiT across three image-text datasets, with Vision Transformer vit, ResNet bit, and MLP-Mixer mixer architectures.”From this paper · §Introduction
- YFCC100M2015 · cited 2דFurthermore, we explore publicly available datasets such as YFCC100m yfcc100m and CC12M cc12m.”From this paper · §Introduction
- Do better ImageNet models transfer bette2018 · cited 2דTransfer learning transfer_learning_survey has been a successful paradigm in computer vision imagenet_transfer_better; bit; instagram_resnext.”From this paper · §Introduction
- T52019 · cited 2דWe consider four possible transformer-based text models transformer—the transformer from ViT-B vit which also resembles that used in CLIP clip, T5-base T5, mT5-base mt5, and the classic BERT-base bert—and whether to init…”From this paper · §Experiments
- Conceptual 12M2021 · cited 2דFurthermore, we explore publicly available datasets such as YFCC100m yfcc100m and CC12M cc12m.”From this paper · §Introduction
- CoAtNet2021 · cited 2דWith the pre-trained model ViT-g/14 vitg, LiT achieves 85.2% zero-shot transfer accuracy on ImageNet, halving the gap between previous best zero-shot transfer results clip; align and supervised fine-tuning results vitg;…”From this paper · §Introduction
Led to
- CoCa2022 · cited 8×, 3 in Method“On the other hand, while many existing methods [32, 33, 35, 30, 36, 37] train model components with multiple stages on various data sources and/or modalities, CoCa is pretrained end-to-end from scratch directly with vari…”From CoCa · §Approach
Abstract
This paper presents contrastive-tuning, a simple method employing contrastive training to align image and text models while still taking advantage of their pre-training. In our empirical study we find that locked pre-trained image models with unlocked text models work best. We call this instance of contrastive-tuning "Locked-image Tuning" (LiT), which just teaches a text model to read out good representations from a pre-trained image model for new tasks. A LiT model gains the capability of zero-shot transfer to new vision tasks, such as image classification or retrieval. The proposed LiT is widely applicable; it works reliably with multiple pre-training methods (supervised and unsupervised) and across diverse architectures (ResNet, Vision Transformers and MLP-Mixer) using three different image-text datasets. With the transformer-based pre-trained ViT-g/14 model, the LiT model achieves 85.2% zero-shot transfer accuracy on the ImageNet test set, and 82.5% on the challenging out-of-distribution ObjectNet test set.