Paper Lineage
Esc
MethodJan 2021arXiv 2101.00529cs.CV

VinVL: Revisiting Visual Representations in Vision-Language Models

Pengchuan Zhang, Xiujun Li, Xiaowei Hu and 5 others

This paper presents a detailed study of improving visual representations for vision language (VL) tasks and develops an improved object detection model to provide object-centric representations of images. Compared to the most widely used \emph{bottom-up and top-down} model \cite{anderson2018bottom}, the new model is bigger, better-designed for VL tasks, and pre-trained on much larger training corpora that combine multiple public annotated object detection datasets.

From the abstract

Built on

12 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (12)

  • Oscar2020 · cited 11×, 4 in Method
    “In this study we pre-train an improved version of Oscar [21], known as Oscar+ models, to learn the joint image-text representations using image tags as anchors for image-text alignment.”
    From this paper · §Oscar+ Pre-training
  • Captions to Visual Concepts2014 · cited 2×, 2 in Method
    “Different from the binary contrastive loss used in Oscar [21], the proposed 3-way Contrastive Loss to effectively optimize the training objectives used for VQA [41] and text-image matching [6]77 7 [6] uses a deep-learnin…”
    From this paper · §Oscar+ Pre-training
  • MS COCO2014 · cited 5×, 1 in Method
    “We build our pre-training corpus based on three types of existing vision and VL datasets: (1) image captioning datasets with human-annotated captions as 𝒘\boldsymbol{w} and machine-generated 55 5 We use the same model t…”
    From this paper · §Oscar+ Pre-training
  • VQA v22016 · cited 3×, 1 in Method
    “We build our pre-training corpus based on three types of existing vision and VL datasets: (1) image captioning datasets with human-annotated captions as 𝒘\boldsymbol{w} and machine-generated 55 5 We use the same model t…”
    From this paper · §Oscar+ Pre-training
Show 8 more
  • “The output of the 𝐕𝐢𝐬𝐢𝐨𝐧Vision module consists of 𝒒{\boldsymbol{q}} and 𝒗\boldsymbol{v}. 𝒒{\boldsymbol{q}} is the semantic representation of the image, such as tags or detected objects, and 𝒗\boldsymbol{v} the…”
    From this paper · §Improving Vision (V) in Vision Language (VL)
  • Grid features for VQA2020 · cited 3×
    “Although [23] shows that the FPN model outperforms the C4 model for object detection, recent studies [14] demonstrate that FPN does not provide more effective region features for VL tasks than C4, which is also confirmed…”
    From this paper · §Improving Vision (V) in Vision Language (VL)
  • Visual Genome2016 · cited 2×
    “Among the aforementioned work, a widely-used object detection (OD) model [2] is trained on the Visual Genome dataset [16].”
    From this paper · §Introduction
  • NLVR22018 · cited 2×
    “We then fine-tune the pre-trained Oscar+ for a wide range of downstream tasks, including VL understanding tasks such as VQA [8], GQA [13], NLVR2 [35], and COCO text-image retrieval [25], and VL generation tasks such as C…”
    From this paper · §Introduction
  • “Compared to the OD model of [2], the new model is better-designed for VL tasks, and is bigger and trained on much larger amounts of data, combining multiple public object detection datasets, including COCO [25], OpenImag…”
    From this paper · §Introduction
  • nocaps2018 · cited 2×
    “To validate the effectiveness of the new OD model, we pre-train a Transformer-based cross-modal fusion model Oscar+ [21] on a public dataset consisting of 8.858.85 million text-image pairs, where the visual representatio…”
    From this paper · §Introduction
  • Unicoder-VL2019 · cited 2×
    “Following [19], we report the top-KK retrieval results on both the 11K and 55K COCO test sets.”
    From this paper · §Adapting to VL Tasks
  • Unified VLP2019 · cited 2×
    “Similar to VLP [21, 45], the self-attention mask is constrained such that a caption token can only attend to the tokens before its position to simulate a uni-directional generation process.”
    From this paper · §Adapting to VL Tasks

Led to

  • ViLT2021 · cited 4×, 2 in Method
    “Following OSCAR (Li et al. 2020b) and VinVL (Zhang et al. 2021), we use the pair method.”
    From ViLT · §Experiments
  • ALBEF2021 · cited 2×
    “The first category focuses on modelling the interactions between image and text features with transformer-based multimodal encoders [10, 11, 12, 13, 1, 14, 15, 2, 3, 16, 8, 17, 18].”
    From ALBEF · §Related Work
  • SimVLM2021 · cited 3×
    “While a variety of approaches have been proposed, a large portion of them require object detection for image region feature regression or tagging as part of the pre-training objectives (Tan & Bansal 2019; Su et al. 2020;…”
    From SimVLM · §Related Work
  • VLMo2021 · cited 4×
    “Pre-training with Transformer [45] backbone networks has substantially advanced the state of the art across natural language processing [34, 10, 28, 22, 11, 36, 1, 7, 8, 4, 5, 6, 31], computer vision [12, 44, 2] and visi…”
    From VLMo · §Related Work
  • METER2021 · cited 7×, 4 in Method
    “Recent works kim2021vilt; xue2021probing; li2021align that tried to adopt vision transformers have not shown satisfactory performance and typically underperform state-of-the-art region-feature-based VLP models (e.g., Vin…”
    From METER · §Introduction
  • Florence2021 · cited 1×, 1 in Method
    “Thus, the object detector has been a de facto tool for image feature extraction, followed by a fusion network for prediction in many works (Anderson et al. 2018; Li et al. 2020; Zhang et al. 2021b; Wang et al. 2020; Fang…”
    From Florence · §Approach
  • VL-BEiT2022 · cited 3×
    “Previous models [24, 35, 21, 46] use an off-the-shelf object detector like Faster R-CNN [30] to obtain image region features.”
    From VL-BEiT · §Related Work
  • BEiT-32022 · cited 4×, 1 in Method
    “In contrast, previous vision-language models [36, 65, 26, 34, 53, 30, 62] usually employ multiple pretraining tasks, such as image-text contrast, image-text matching, and word-patch/region alignment.”
    From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
Abstract

This paper presents a detailed study of improving visual representations for vision language (VL) tasks and develops an improved object detection model to provide object-centric representations of images. Compared to the most widely used \emph{bottom-up and top-down} model \cite{anderson2018bottom}, the new model is bigger, better-designed for VL tasks, and pre-trained on much larger training corpora that combine multiple public annotated object detection datasets. Therefore, it can generate representations of a richer collection of visual objects and concepts. While previous VL research focuses mainly on improving the vision-language fusion model and leaves the object detection model improvement untouched, we show that visual features matter significantly in VL models. In our experiments we feed the visual features generated by the new object detection model into a Transformer-based VL fusion model \oscar \cite{li2020oscar}, and utilize an improved approach \short\ to pre-train the VL model and fine-tune it on a wide range of downstream VL tasks. Our results show that the new visual features significantly improve the performance across all VL tasks, creating new state-of-the-art results on seven public benchmarks. We will release the new object detection model to public.