Paper Lineage
Esc
MethodMay 2022arXiv 2205.01917cs.CV

CoCa: Contrastive Captioners are Image-Text Foundation Models

Jiahui Yu, Zirui Wang, Vijay Vasudevan and 3 others

Exploring large-scale pretrained foundation models is of significant interest in computer vision because these models can be quickly transferred to many downstream tasks. This paper presents Contrastive Captioner (CoCa), a minimalist design to pretrain an image-text encoder-decoder foundation model jointly with contrastive loss and captioning loss, thereby subsuming model capabilities from contrastive approaches like CLIP and generative methods like SimVLM.

From the abstract

Built on

16 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (16)

  • ALIGN2021 · cited 13×, 3 in Method
    “Following ALIGN [13], we pretrain with image resolution of 288×\times288 and patch size 18×\times18, resulting in a total of 256 image tokens.”
    From this paper · §Approach
  • LiT2021 · cited 8×, 3 in Method
    “On the other hand, while many existing methods [32, 33, 35, 30, 36, 37] train model components with multiple stages on various data sources and/or modalities, CoCa is pretrained end-to-end from scratch directly with vari…”
    From this paper · §Approach
  • CLIP2021 · cited 13×, 2 in Method
    “Following previous practices [12, 32], “zero-shot” here is different from classical zero-shot learning in that during pretraining, the model may see relevant supervised information, but no supervised examples are used du…”
    From this paper · §Approach
  • ALBEF2021 · cited 4×, 2 in Method
    “Since unidirectional language models are trained with causal masking on complete sentences, the decoder can efficiently generate outputs for both contrastive and generative losses with a single forward propagation (compa…”
    From this paper · §Approach
Show 12 more
  • ResNet2015 · cited 2×, 2 in Method
    “Following a standard encoder-decoder architecture, the image encoder provides latent encoded features (e.g., using a Vision Transformer [39] or ConvNets [40]) and the text decoder learns to maximize the conditional likel…”
    From this paper · §Approach
  • Set Transformer2018 · cited 2×, 2 in Method
    “As discussed in the previous section, CoCa adopts task-specific attentional pooling [42] (pooler for brevity) to customize visual representations for different types downstream tasks while sharing the backbone encoder.”
    From this paper · §Approach
  • ViT2020 · cited 2×, 2 in Method
    “Following a standard encoder-decoder architecture, the image encoder provides latent encoded features (e.g., using a Vision Transformer [39] or ConvNets [40]) and the text decoder learns to maximize the conditional likel…”
    From this paper · §Approach
  • SimVLM2021 · cited 9×, 1 in Method
    “Unlike other fusion-based foundation methods [16, 35, 17], CoCa is naturally applicable to crossmodal alignment tasks since it generates aligned image and text unimodal embeddings.”
    From this paper · §Experiments
  • Instagram hashtag pre-training2018 · cited 2×, 1 in Method
    “The classic single-encoder approach pretrains a visual encoder through image classification on a large crowd-sourced image annotation dataset (e.g., ImageNet [9], Instagram [20] or JFT [21]), where the vocabulary of anno…”
    From this paper · §Approach
  • VLMo2021 · cited 2×, 1 in Method
    “On the other hand, while many existing methods [32, 33, 35, 30, 36, 37] train model components with multiple stages on various data sources and/or modalities, CoCa is pretrained end-to-end from scratch directly with vari…”
    From this paper · §Approach
  • MAE2021 · cited 2×, 1 in Method
    “MAE [23] and SimMIM [24] remove the need for an image tokenizer and directly use a light-weight decoder or projection layer to regress pixel values.”
    From this paper · §Related Work
  • BLIP2022 · cited 2×, 1 in Method
    “On the other hand, while many existing methods [32, 33, 35, 30, 36, 37] train model components with multiple stages on various data sources and/or modalities, CoCa is pretrained end-to-end from scratch directly with vari…”
    From this paper · §Approach
  • Florence2021 · cited 6×
    “However, these models rely heavily on image annotations as labeled vectors and do not bake in knowledge of free-form human natural language, hindering their application to downstream tasks that involving both vision and…”
    From this paper · §Introduction
  • COCO Captions2015 · cited 2×
    “We evaluate CoCa on the two standard image-text retrieval benchmarks: MSCOCO [63] and Flickr30K [62].”
    From this paper · §Experiments
  • BERT2018 · cited 2×
    “Pretraining ConvNets [18] or Transformers [19] on large-scale annotated data such as ImageNet [6, 7, 8], Instagram [20] or JFT [21] has become a popular strategy towards solving visual recognition problems including clas…”
    From this paper · §Related Work
  • T52019 · cited 2×
    “However, these models rely heavily on image annotations as labeled vectors and do not bake in knowledge of free-form human natural language, hindering their application to downstream tasks that involving both vision and…”
    From this paper · §Introduction

Led to

  • BEiT-32022 · cited 6×, 4 in Method
    “In contrast, previous vision-language models [36, 65, 26, 34, 53, 30, 62] usually employ multiple pretraining tasks, such as image-text contrast, image-text matching, and word-patch/region alignment.”
    From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
Abstract

Exploring large-scale pretrained foundation models is of significant interest in computer vision because these models can be quickly transferred to many downstream tasks. This paper presents Contrastive Captioner (CoCa), a minimalist design to pretrain an image-text encoder-decoder foundation model jointly with contrastive loss and captioning loss, thereby subsuming model capabilities from contrastive approaches like CLIP and generative methods like SimVLM. In contrast to standard encoder-decoder transformers where all decoder layers attend to encoder outputs, CoCa omits cross-attention in the first half of decoder layers to encode unimodal text representations, and cascades the remaining decoder layers which cross-attend to the image encoder for multimodal image-text representations. We apply a contrastive loss between unimodal image and text embeddings, in addition to a captioning loss on the multimodal decoder outputs which predicts text tokens autoregressively. By sharing the same computational graph, the two training objectives are computed efficiently with minimal overhead. CoCa is pretrained end-to-end and from scratch on both web-scale alt-text data and annotated images by treating all labels simply as text, seamlessly unifying natural language supervision for representation learning. Empirically, CoCa achieves state-of-the-art performance with zero-shot transfer or minimal task-specific adaptation on a broad range of downstream tasks, spanning visual recognition (ImageNet, Kinetics-400/600/700, Moments-in-Time), crossmodal retrieval (MSCOCO, Flickr30K, MSR-VTT), multimodal understanding (VQA, SNLI-VE, NLVR2), and image captioning (MSCOCO, NoCaps). Notably on ImageNet classification, CoCa obtains 86.3% zero-shot top-1 accuracy, 90.6% with a frozen encoder and learned classification head, and new state-of-the-art 91.0% top-1 accuracy on ImageNet with a finetuned encoder.