Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts
The availability of large-scale image captioning and visual question answering datasets has contributed significantly to recent successes in vision-and-language pre-training. , image caption generation), which limit the resulting dataset scale and diversity.
Also cited · not yet reviewed (28)
- ViLBERT2019 · cited 12×, 5 in Method“The large scale nature and the high degree of textual and visual diversity make this dataset particularly suited to V+L pre-training [55, 21, 74, 88, 48, 56].”From this paper · §Vision-and-Language Pre-Training Data
- Decoupled box proposals captioning2019 · cited 6×, 4 in Method“We select top-16 box proposals and featurize each of them with Graph-RISE, similar to [19].”From this paper · §Evaluating Vision-and-Language Pre-Training Data
- Localized Narratives2019 · cited 3×, 3 in Method“In addition, besides CC3M and CC12M, we also explore using the Open Images Localized Narratives dataset (LocNar) [64], as an alternative “in-domain” (from a visual standpoint) pre-training data source.”From this paper · §Evaluating Vision-and-Language Pre-Training Data
- Unified VLP2019 · cited 7×, 2 in Method“The large scale nature and the high degree of textual and visual diversity make this dataset particularly suited to V+L pre-training [55, 21, 74, 88, 48, 56].”From this paper · §Vision-and-Language Pre-Training Data
Show 24 more
- MS COCO2014 · cited 4×, 2 in Method“Unlike in the standard image captioning setting, nocaps’s distributions of images during training (COCO Captions) and evaluation (Open Images) are different: the Open Images dataset [45, 42] covers one order of magnitude…”From this paper · §Evaluating Vision-and-Language Pre-Training Data
- The Open Images Dataset V42018 · cited 2×, 2 in Method“Unlike in the standard image captioning setting, nocaps’s distributions of images during training (COCO Captions) and evaluation (Open Images) are different: the Open Images dataset [45, 42] covers one order of magnitude…”From this paper · §Evaluating Vision-and-Language Pre-Training Data
- UNITER2019 · cited 7×, 1 in Method“The large scale nature and the high degree of textual and visual diversity make this dataset particularly suited to V+L pre-training [55, 21, 74, 88, 48, 56].”From this paper · §Vision-and-Language Pre-Training Data
- Unicoder-VL2019 · cited 6×, 1 in Method“The large scale nature and the high degree of textual and visual diversity make this dataset particularly suited to V+L pre-training [55, 21, 74, 88, 48, 56].”From this paper · §Vision-and-Language Pre-Training Data
- VL-BERT2019 · cited 6×, 1 in Method“The large scale nature and the high degree of textual and visual diversity make this dataset particularly suited to V+L pre-training [55, 21, 74, 88, 48, 56].”From this paper · §Vision-and-Language Pre-Training Data
- 12-in-12019 · cited 6×, 1 in Method“The large scale nature and the high degree of textual and visual diversity make this dataset particularly suited to V+L pre-training [55, 21, 74, 88, 48, 56].”From this paper · §Vision-and-Language Pre-Training Data
- Visual Genome2016 · cited 5×, 1 in Method“We train a Faster-RCNN [68] on Visual Genome [43], with a ResNet101 [29] backbone trained on JFT [30] and fine-tuned on ImageNet [69].”From this paper · §Evaluating Vision-and-Language Pre-Training Data
- BERT2018 · cited 5×, 1 in Method“For instance, Zhou et al. [88] adapt BERT [25] to generate text.”From this paper · §Evaluating Vision-and-Language Pre-Training Data
- LXMERT2019 · cited 5×, 1 in Method“To train the model’s parameters, we use a contrastive softmax loss, for which the original image-text pairs are used as positive examples, while all other image-text pairs in the mini-batch are used as negative examples…”From this paper · §Evaluating Vision-and-Language Pre-Training Data
- nocaps2018 · cited 4×, 1 in Method“Table 4 compares our best model (ic pre-trained on CC3M+CC12M) to existing state-of-the-art results on nocaps, and show that ours achieves state-of-the-art performance on CIDEr, outperforming a concurrent work [32] that…”From this paper · §Experimental Results
- Oscar2020 · cited 3×, 1 in Method“Inspired by [50], we obtain up to 16 image tags from the Google Cloud Vision APIs, and treat them as text inputs to our model.”From this paper · §Evaluating Vision-and-Language Pre-Training Data
- Transformer2017 · cited 2×, 1 in Method“For ic-based pre-training and downstream tasks, we follow the state-of-the-art architecture that heavily rely on self-attention [78] or similar mechanisms [70, 85, 19, 33, 23].”From this paper · §Evaluating Vision-and-Language Pre-Training Data
- CIDEr2014 · cited 1×, 1 in Method“To measure the performance on image caption generation, we consider the standard metrics BLEU-1,4 [62], ROUGE-L [51], METEOR [10], CIDEr-D [79], and SPICE [4].”From this paper · §Evaluating Vision-and-Language Pre-Training Data
- Knowledge Distillation2015 · cited 1×, 1 in Method“We train a Faster-RCNN [68] on Visual Genome [43], with a ResNet101 [29] backbone trained on JFT [30] and fine-tuned on ImageNet [69].”From this paper · §Evaluating Vision-and-Language Pre-Training Data
- Flickr30k Entities2015 · cited 1×, 1 in Method“The Flickr30K dataset [63] consists of 31,000 images from Flickr, each associated with five captions.”From this paper · §Evaluating Vision-and-Language Pre-Training Data
- Faster R-CNN2015 · cited 1×, 1 in Method“We train a Faster-RCNN [68] on Visual Genome [43], with a ResNet101 [29] backbone trained on JFT [30] and fine-tuned on ImageNet [69].”From this paper · §Evaluating Vision-and-Language Pre-Training Data
- ResNet2015 · cited 1×, 1 in Method“We train a Faster-RCNN [68] on Visual Genome [43], with a ResNet101 [29] backbone trained on JFT [30] and fine-tuned on ImageNet [69].”From this paper · §Evaluating Vision-and-Language Pre-Training Data
- Bottom-Up Top-Down attention2017 · cited 1×, 1 in Method“We use Graph-RISE [37, 38] to featurize the entire image.”From this paper · §Evaluating Vision-and-Language Pre-Training Data
- Meshed-Memory Transformer2019 · cited 1×, 1 in Method“For ic-based pre-training and downstream tasks, we follow the state-of-the-art architecture that heavily rely on self-attention [78] or similar mechanisms [70, 85, 19, 33, 23].”From this paper · §Evaluating Vision-and-Language Pre-Training Data
- COCO Captions2015 · cited 4דSmaller but less noisy SBU Captions [61] (1̃M) and COCO Captions [20] (106K) datasets are also of high interest.”From this paper · §Related Work
- VisualBERT2019 · cited 3דBased directly upon BERT, V+L pre-training research has largely been focused on V+L understanding [55, 49, 21, 77, 3, 74, 48, 56], with classification or regression tasks that do not involve generation.”From this paper · §Related Work
- Constrained beam search captioning2016 · cited 2דAddressing long-tail distributions of visual concepts is an important component of V+L systems that generalize, as long and free-form texts exhibit a large number of compositional, fine-grained categories [89, 54, 19].”From this paper · §Related Work
- VQA v22016 · cited 2דSecond, many of the small-sized datasets share the same, limited visual domain; COCO-Captions [20], Visual Genome [44], and VQA2 [27] are (mostly) based on several hundreds thousand of COCO images [52].”From this paper · §Introduction
- T52019 · cited 2דOn one hand, this is due to advances in architectures and modeling that are mainly inspired by BERT and similar models in natural language understanding and generation [25, 53, 82, 46, 26, 66].”From this paper · §Introduction
Led to
- ALBEF2021 · cited 1×, 1 in Method“To show that our method is scalable with larger-scale web data, we also include the much noisier Conceptual 12M dataset [43], increasing the total number of images to 14.1M 33 3 some urls provided by the web datasets hav…”From ALBEF · §ALBEF Pre-training
- LiT2021 · cited 2דFurthermore, we explore publicly available datasets such as YFCC100m yfcc100m and CC12M cc12m.”From LiT · §Introduction
- ViTCAP2021 · cited 2×, 1 in Method“CC-1212M is the model trained with 1212M image-caption pairs changpinyo2021conceptual.”From ViTCAP · §Experiment
- BLIP2022 · cited 3דDue to the prohibitive expense of acquiring human-annotated texts, most methods (Chen et al. 2020; Li et al. 2020; Li et al. 2021a; Wang et al. 2021; Radford et al. 2021) use image and alt-text pairs crawled from the web…”From BLIP · §Related Work
- BEiT-32022 · cited 1×, 1 in Method“For multimodal data, there are about 1515M images and 2121M image-text pairs collected from five public datasets: Conceptual 12M (CC12M) [11], Conceptual Captions (CC3M) [46], SBU Captions (SBU) [39], COCO [31] and Visua…”From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
- BLIP-22023 · cited 1×, 1 in Method“We use the same pre-training dataset as BLIP with 129M images in total, including COCO (Lin et al. 2014), Visual Genome (Krishna et al. 2017), CC3M (Sharma et al. 2018), CC12M (Changpinyo et al. 2021), SBU (Ordonez et al…”From BLIP-2 · §Method
Abstract
The availability of large-scale image captioning and visual question answering datasets has contributed significantly to recent successes in vision-and-language pre-training. However, these datasets are often collected with overrestrictive requirements inherited from their original target tasks (e.g., image caption generation), which limit the resulting dataset scale and diversity. We take a step further in pushing the limits of vision-and-language pre-training data by relaxing the data collection pipeline used in Conceptual Captions 3M (CC3M) [Sharma et al. 2018] and introduce the Conceptual 12M (CC12M), a dataset with 12 million image-text pairs specifically meant to be used for vision-and-language pre-training. We perform an analysis of this dataset and benchmark its effectiveness against CC3M on multiple downstream tasks with an emphasis on long-tail visual recognition. Our results clearly illustrate the benefit of scaling up pre-training data for vision-and-language tasks, as indicated by the new state-of-the-art results on both the nocaps and Conceptual Captions benchmarks.