Contrastive image–text models
Dual encoders that align images and text in one embedding space, enabling zero-shot transfer.
| Concept | Introduced by | Years | Papers |
|---|---|---|---|
| Joint image–text embedding Map images and sentences into one space so they can be retrieved by similarity. | Deep Fragment Embeddings | 2014 | 0 |
| Contrastive image–text learning Pull matching image and text embeddings together and push mismatched pairs apart. | ConVIRT | 2020–2021 | 1 |
| Locked image tuning Keep a pretrained image tower frozen and train only the text side contrastively. | LiT | 2021–2023 | 2 |
| Scaling on raw alt-text Skip expensive cleaning: a billion noisy alt-text pairs beat curated datasets. | ALIGN | 2021–2022 | 12 |
| Vision foundation model One large image–text model adapted to classification, retrieval, detection, video and more. | Florence | 2021–2022 | 5 |
| Zero-shot transfer via text prompts Classify into unseen classes by comparing an image with text descriptions of them. | CLIP | 2021–2023 | 21 |
| Contrastive + captioning in one model One image–text encoder–decoder trained with both a contrastive and a captioning loss. | CoCa | 2022 | 0 |
| Unified image–text–label contrastive Treat class labels and captions alike in one contrastive space. | UniCL | 2022 | 0 |