Unified Contrastive Learning in Image-Text-Label Space
Visual recognition is recently learned via either supervised learning on human-annotated image-label data or language-image contrastive learning with webly-crawled image-text pairs. While supervised learning may result in a more discriminative representation, language-image pretraining shows unprecedented zero-shot recognition capability, largely due to the different properties of data sources and learning objectives.
Also cited · not yet reviewed (7)
- CLIP2021 · cited 14×, 1 in Method“Typically, this can be tackled via either supervised learning on human-annotated image-label pairs deng2009imagenet or contrastive learning on webly-crawed image-text pairs radford2021learning; jia2021scaling.”From this paper · §Introduction
- ALIGN2021 · cited 5דTypically, this can be tackled via either supervised learning on human-annotated image-label pairs deng2009imagenet or contrastive learning on webly-crawed image-text pairs radford2021learning; jia2021scaling.”From this paper · §Introduction
- ResNet2015 · cited 3דWhen fueled with clean and large-scale human-annotated image-label data, e.g., ImageNet deng2009imagenet, supervised learning can attain decent visual recognition capacities over the given categories krizhevsky2012imagen…”From this paper · §Introduction
- Inception v32015 · cited 2דWhen fueled with clean and large-scale human-annotated image-label data, e.g., ImageNet deng2009imagenet, supervised learning can attain decent visual recognition capacities over the given categories krizhevsky2012imagen…”From this paper · §Introduction
Show 3 more
- SimCLR2020 · cited 2דContrastive learning has laid the foundation for the best performing SSL models tian2019contrastive; henaff2019data; chen2020simple; he2020momentum; caron2020unsupervised; tian2020makes; chen2021empirical.”From this paper · §Related works
- Swin2021 · cited 2דWith this goal, numerous works have pushed the image recognition performance from different directions, such as data scale from MNIST lecun1989handwritten to ImageNet-1K deng2009imagenet, model architectures from convolu…”From this paper · §Related works
- Florence2021 · cited 2דFinally, we scaled up UniCL to billions of image-text-label data in Florence yuan2021florence and demonstrated its superiority over CLIP radford2021learning and ALIGN jia2021scaling across dozens of benchmarks.”From this paper · §Introduction
Led to
Nothing yet. No later paper in this dataset cites it strongly. Suggest a paper that builds on UniCL →
Abstract
Visual recognition is recently learned via either supervised learning on human-annotated image-label data or language-image contrastive learning with webly-crawled image-text pairs. While supervised learning may result in a more discriminative representation, language-image pretraining shows unprecedented zero-shot recognition capability, largely due to the different properties of data sources and learning objectives. In this work, we introduce a new formulation by combining the two data sources into a common image-text-label space. In this space, we propose a new learning paradigm, called Unified Contrastive Learning (UniCL) with a single learning objective to seamlessly prompt the synergy of two data types. Extensive experiments show that our UniCL is an effective way of learning semantically rich yet discriminative representations, universally for image recognition in zero-shot, linear-probe, fully finetuning and transfer learning scenarios. Particularly, it attains gains up to 9.2% and 14.5% in average on zero-shot recognition benchmarks over the language-image contrastive learning and supervised learning methods, respectively. In linear probe setting, it also boosts the performance over the two methods by 7.3% and 3.4%, respectively. Our study also indicates that UniCL stand-alone is a good learner on pure image-label data, rivaling the supervised learning methods across three image classification datasets and two types of vision backbones, ResNet and Swin Transformer. Code is available at https://github.com/microsoft/UniCL.