Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
Large-scale pre-training methods of learning cross-modal representations on image-text pairs are becoming popular for vision-language tasks. While existing methods simply concatenate image region features and text features as input to the model to be pre-trained and use self-attention to learn image-text semantic alignments in a brute force manner, in this paper, we propose a new learning method Oscar (Object-Semantics Aligned Pre-training), which uses object tags detected in images as anchor points to significantly ease the learning of alignments.
Also cited · not yet reviewed (4)
- UNITER2019 · cited 5דThe existing methods [37, 38, 22, 5, 46, 35, 19, 10] employ BERT-like objectives [6] to learn cross-modal representations from a concatenated-sequence of visual region features and language token embeddings.”From this paper · §Related Work
- Unicoder-VL2019 · cited 3דFollowing [19], we report the top-KK retrieval results on both the 11K and 55K COCO test sets.”From this paper · §Adapting to V+L Tasks
- The Open Images Dataset V42018 · cited 2דTo study the impact of different object tag sets in pre-trained models, we pre-train two variants: OscarVG^{VG} and OscarOI^{OI} utilizes object tags produced by the object detector trained on the visual genome (VG) data…”From this paper · §Experimental Results & Analysis
- VL-BERT2019 · cited 2דThe existing methods [37, 38, 22, 5, 46, 35, 19, 10] employ BERT-like objectives [6] to learn cross-modal representations from a concatenated-sequence of visual region features and language token embeddings.”From this paper · §Related Work
Led to
- VinVL2021 · cited 11×, 4 in Method“In this study we pre-train an improved version of Oscar [21], known as Oscar+ models, to learn the joint image-text representations using image tags as anchors for image-text alignment.”From VinVL · §Oscar+ Pre-training
- VL-T52021 · cited 4דFollowing this success, image+text pretraining models (Lu et al. 2019; Tan & Bansal 2019; Chen et al. 2020; Huang et al. 2020; Li et al. 2020b; Cho et al. 2020; Radford et al. 2021; Zhang et al. 2021) and video+text pret…”From VL-T5 · §Related Works
- ViLT2021 · cited 2דFollowing OSCAR (Li et al. 2020b) and VinVL (Zhang et al. 2021), we use the pair method.”From ViLT · §Experiments
- ALIGN2021 · cited 2דPre-training has also become the de-facto approach in vision-language modeling (Lu et al. 2019; Chen et al. 2020c; Li et al. 2020).”From ALIGN · §Introduction
- Conceptual 12M2021 · cited 3×, 1 in Method“Inspired by [50], we obtain up to 16 image tags from the Google Cloud Vision APIs, and treat them as text inputs to our model.”From Conceptual 12M · §Evaluating Vision-and-Language Pre-Training Data
- ALBEF2021 · cited 5דWe evaluate ALBEF on the Flickr30K [49] and COCO benchmarks, and fine-tune the pre-trained model using the training samples from each dataset.”From ALBEF · §Downstream V+L Tasks
- SimVLM2021 · cited 4דWhile a variety of approaches have been proposed, a large portion of them require object detection for image region feature regression or tagging as part of the pre-training objectives (Tan & Bansal 2019; Su et al. 2020;…”From SimVLM · §Related Work
- VLMo2021 · cited 3דFollowing OSCAR [26] and VinVL [49], we convert the triplet input to two image-text pairs, each containing the text description and one image.”From VLMo · §Experiments
- METER2021 · cited 4×, 2 in Method“Vision-and-language pre-training (VLP) has now become the de facto practice to tackle these tasks tan-bansal-2019-lxmert; li2019visualbert; lu2019vilbert; su2019vl; chen2020uniter; li2020oscar.”From METER · §Introduction
- Florence2021 · cited 1×, 1 in Method“Thus, the object detector has been a de facto tool for image feature extraction, followed by a fusion network for prediction in many works (Anderson et al. 2018; Li et al. 2020; Zhang et al. 2021b; Wang et al. 2020; Fang…”From Florence · §Approach
- ViTCAP2021 · cited 10דRecent studies have witnessed its great development which are primarily reflected in the aspects of more advanced cross-modal fusion architectures you2016image; vinyals2015show; xu2015show; rennie2017self; yang2019auto;…”From ViTCAP · §Introduction
- BLIP2022 · cited 2דDue to the prohibitive expense of acquiring human-annotated texts, most methods (Chen et al. 2020; Li et al. 2020; Li et al. 2021a; Wang et al. 2021; Radford et al. 2021) use image and alt-text pairs crawled from the web…”From BLIP · §Related Work
- Simple end-to-end captioning2022 · cited 5×, 3 in Method“We follow the suggestion of Li et al. 2020 to evaluate our models with the validation set, containing 4.5k images and 10 captions per image.”From Simple end-to-end captioning · §. Experiment Setup
- VL-BEiT2022 · cited 3דPrevious models [24, 35, 21, 46] use an off-the-shelf object detector like Faster R-CNN [30] to obtain image region features.”From VL-BEiT · §Related Work
- BEiT-32022 · cited 1×, 1 in Method“In contrast, previous vision-language models [36, 65, 26, 34, 53, 30, 62] usually employ multiple pretraining tasks, such as image-text contrast, image-text matching, and word-patch/region alignment.”From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
Abstract
Large-scale pre-training methods of learning cross-modal representations on image-text pairs are becoming popular for vision-language tasks. While existing methods simply concatenate image region features and text features as input to the model to be pre-trained and use self-attention to learn image-text semantic alignments in a brute force manner, in this paper, we propose a new learning method Oscar (Object-Semantics Aligned Pre-training), which uses object tags detected in images as anchor points to significantly ease the learning of alignments. Our method is motivated by the observation that the salient objects in an image can be accurately detected, and are often mentioned in the paired text. We pre-train an Oscar model on the public corpus of 6.5 million text-image pairs, and fine-tune it on downstream tasks, creating new state-of-the-arts on six well-established vision-language understanding and generation tasks.