Paper Lineage
Esc
MethodFeb 2021arXiv 2102.05918cs.CV

Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Chao Jia, Yinfei Yang, Ye Xia and 7 others

Pre-trained representations are becoming crucial for many NLP and perception tasks. While representation learning in NLP has transitioned to training on raw text without human annotations, visual and vision-language representations still rely heavily on curated training datasets that are expensive or require expert knowledge.

From the abstract

Built on

13 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (13)

  • Flickr30k Entities2015 · cited 2×, 2 in Method
    “Two benchmark datasets are considered: Flickr30K (Plummer et al. 2015) and MSCOCO (Chen et al. 2015).”
    From this paper · §Pre-training and Task Transfer
  • BiT2019 · cited 5×, 1 in Method
    “Following Kolesnikov et al. 2020, we also evaluate the robustness of our model on Visual Task Adaptation Benchmark (VTAB) (Zhai et al. 2019) which consists of 19 diverse (covering subgroups of natural, specialized and st…”
    From this paper · §Pre-training and Task Transfer
  • COCO Captions2015 · cited 1×, 1 in Method
    “Two benchmark datasets are considered: Flickr30K (Plummer et al. 2015) and MSCOCO (Chen et al. 2015).”
    From this paper · §Pre-training and Task Transfer
  • ImageNetV22019 · cited 1×, 1 in Method
    “We first apply zero-shot transfer of ALIGN to visual classification tasks on ImageNet ILSVRC-2012 benchmark (Deng et al. 2009) and its variants including ImageNet-R(endition) (Hendrycks et al. 2020) (non-natural images s…”
    From this paper · §Pre-training and Task Transfer
Show 9 more
  • UNITER2019 · cited 3×
    “Recently more advanced models emerge with cross-modal attention layers (Liu et al. 2019a; Lu et al. 2019; Chen et al. 2020c; Huang et al. 2020b) and show superior performance in image-text matching tasks.”
    From this paper · §Related Work
  • “Joulin et al. 2015; Li et al. 2017; Desai & Johnson 2020; Sariyildiz et al. 2020; Zhang et al. 2020 show that a good visual representation can be learned by predicting the captions from images, which inspires our work.”
    From this paper · §Related Work
  • Visual N-Grams2016 · cited 2×
    “Joulin et al. 2015; Li et al. 2017; Desai & Johnson 2020; Sariyildiz et al. 2020; Zhang et al. 2020 show that a good visual representation can be learned by predicting the captions from images, which inspires our work.”
    From this paper · §Related Work
  • ViLBERT2019 · cited 2×
    “Recently more advanced models emerge with cross-modal attention layers (Liu et al. 2019a; Lu et al. 2019; Chen et al. 2020c; Huang et al. 2020b) and show superior performance in image-text matching tasks.”
    From this paper · §Related Work
  • SimCLR2020 · cited 2×
    “Recently, self-supervised (Chen et al. 2020b; Tian et al. 2020; He et al. 2020; Misra & Maaten 2020; Li et al. 2021; Grill et al. 2020; Caron et al. 2020) and semi-supervised learning (Yalniz et al. 2019; Xie et al. 2020…”
    From this paper · §Related Work
  • Oscar2020 · cited 2×
    “Pre-training has also become the de-facto approach in vision-language modeling (Lu et al. 2019; Chen et al. 2020c; Li et al. 2020).”
    From this paper · §Introduction
  • VirTex2020 · cited 2×
    “Joulin et al. 2015; Li et al. 2017; Desai & Johnson 2020; Sariyildiz et al. 2020; Zhang et al. 2020 show that a good visual representation can be learned by predicting the captions from images, which inspires our work.”
    From this paper · §Related Work
  • ViT2020 · cited 2×
    “High-quality visual representations for classification or retrieval are usually pre-trained on large-scale labeled datasets (Mahajan et al. 2018; Kolesnikov et al. 2020; Dosovitskiy et al. 2021; Juan et al. 2020).”
    From this paper · §Related Work
  • CLIP2021 · cited 2×
    “In the zero-shot setting, ALIGN gets more than 7% improvement in image retrieval task compared to the previous SOTA, CLIP (Radford et al. 2021).”
    From this paper · §Experiments and Results

Led to

  • ALBEF2021 · cited 4×
    “The first category focuses on modelling the interactions between image and text features with transformer-based multimodal encoders [10, 11, 12, 13, 1, 14, 15, 2, 3, 16, 8, 17, 18].”
    From ALBEF · §Related Work
  • SimVLM2021 · cited 6×
    “Motivated by recent works (Radford et al. 2021; Ramesh et al. 2021; Jia et al. 2021; Tsimpoukelli et al. 2021) that illustrate zero-shot learning in certain image-text tasks, we train our model using large-scale weakly l…”
    From SimVLM · §Related Work
  • LAION-400M2021 · cited 2×
    “Multi-modal language-vision models demonstrated recently strong transfer capability to novel datasets in absense of per-sample labels [1, 2, 3].”
    From LAION-400M · §Introduction
  • VLMo2021 · cited 3×
    “Pre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”
    From VLMo · §Related Work
  • LiT2021 · cited 14×, 3 in Method
    “Recently it was demonstrated that web-sourced paired image-text data can be used to pre-train strong models for zero-shot transfer clip; align.”
    From LiT · §Introduction
  • Florence2021 · cited 8×, 1 in Method
    “To improve data quality, we performed rigorous data filtering, similar to ALIGN (Jia et al. 2021), including a simple hash-based near-duplicate image removal, small-size image removal, image-text relevance, etc.”
    From Florence · §Approach
  • UniCL2022 · cited 5×
    “Typically, this can be tackled via either supervised learning on human-annotated image-label pairs deng2009imagenet or contrastive learning on webly-crawed image-text pairs radford2021learning; jia2021scaling.”
    From UniCL · §Introduction
  • Flamingo2022 · cited 4×, 1 in Method
    “Another family of vision-language models is based on contrastive learning [2, 85, 50, 146, 82, 74, 5, 140, 57, 138, 49].”
    From Flamingo · §Related work
  • CoCa2022 · cited 13×, 3 in Method
    “Following ALIGN [13], we pretrain with image resolution of 288×\times288 and patch size 18×\times18, resulting in a total of 256 image tokens.”
    From CoCa · §Approach
  • VL-BEiT2022 · cited 2×
    “Dual-encoder model [29, 14] consists of an image encoder and a text encoder.”
    From VL-BEiT · §Related Work
  • BEiT-32022 · cited 2×, 2 in Method
    “In comparison, contrastive-based models [43, 23, 60, 62] usually need a very large batch size11 1 For example, CoCa [62] uses 6565k batch size, CLIP [43] uses 3232k batch size, and Florence [60] uses 2424k batch size.”
    From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
Abstract

Pre-trained representations are becoming crucial for many NLP and perception tasks. While representation learning in NLP has transitioned to training on raw text without human annotations, visual and vision-language representations still rely heavily on curated training datasets that are expensive or require expert knowledge. For vision applications, representations are mostly learned using datasets with explicit class labels such as ImageNet or OpenImages. For vision-language, popular datasets like Conceptual Captions, MSCOCO, or CLIP all involve a non-trivial data collection (and cleaning) process. This costly curation process limits the size of datasets and hence hinders the scaling of trained models. In this paper, we leverage a noisy dataset of over one billion image alt-text pairs, obtained without expensive filtering or post-processing steps in the Conceptual Captions dataset. A simple dual-encoder architecture learns to align visual and language representations of the image and text pairs using a contrastive loss. We show that the scale of our corpus can make up for its noise and leads to state-of-the-art representations even with such a simple learning scheme. Our visual representation achieves strong performance when transferred to classification tasks such as ImageNet and VTAB. The aligned visual and language representations enables zero-shot image classification and also set new state-of-the-art results on Flickr30K and MSCOCO image-text retrieval benchmarks, even when compared with more sophisticated cross-attention models. The representations also enable cross-modality search with complex text and text + image queries.