Paper Lineage
Esc

Contrastive image–text models

Dual encoders that align images and text in one embedding space, enabling zero-shot transfer.

Concepts in Contrastive image–text models
ConceptIntroduced byYearsPapers
Joint image–text embedding
Map images and sentences into one space so they can be retrieved by similarity.
Deep Fragment Embeddings20140
Contrastive image–text learning
Pull matching image and text embeddings together and push mismatched pairs apart.
ConVIRT2020–20211
Locked image tuning
Keep a pretrained image tower frozen and train only the text side contrastively.
LiT2021–20232
Scaling on raw alt-text
Skip expensive cleaning: a billion noisy alt-text pairs beat curated datasets.
ALIGN2021–202212
Vision foundation model
One large image–text model adapted to classification, retrieval, detection, video and more.
Florence2021–20225
Zero-shot transfer via text prompts
Classify into unseen classes by comparing an image with text descriptions of them.
CLIP2021–202321
Contrastive + captioning in one model
One image–text encoder–decoder trained with both a contrastive and a captioning loss.
CoCa20220
Unified image–text–label contrastive
Treat class labels and captions alike in one contrastive space.
UniCL20220