Paper Lineage
Esc
MethodDec 2019arXiv 1912.02315cs.CV

12-in-1: Multi-Task Vision and Language Representation Learning

Jiasen Lu, Vedanuj Goswami, Marcus Rohrbach and 2 others

Much of vision-and-language research focuses on a small but diverse set of independent tasks and supporting datasets often studied in isolation; however, the visually-grounded language understanding skills required for success at these tasks overlap significantly. In this work, we investigate these relationships between vision-and-language tasks by developing a large-scale, multi-task training regime.

From the abstract

Built on

14 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (14)

  • ViLBERT2019 · cited 9×, 4 in Method
    “The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”
    From this paper · §Introduction
  • Unicoder-VL2019 · cited 6×, 2 in Method
    “The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”
    From this paper · §Introduction
  • decaNLP2018 · cited 5×, 2 in Method
    “Inspired by prior multi-task literature bengio2009curriculum mccann2018natural, we experimented with both curriculum and anti-curriculum strategies based on task difficulty.”
    From this paper · §Approach
  • Bottom-Up Top-Down attention2017 · cited 2×, 2 in Method
    “At the interface level, ViLBERT takes as input an image II and text segment QQ represented as the sequence {\{IMG,v1,…,v𝒯,,v_{1},\dots,v_{T}, CLS, w1,…,wT,w_{1},\dots,w_{T}, SEP}\} where {vi}i=1𝒯\{v_{i}\}_{i=1}^{T} are…”
    From this paper · §Approach
Show 10 more
  • Multi-task hierarchical VL2018 · cited 6×, 1 in Method
    “Different from previous observation mccann2018natural; nguyen2019multi, we found that using no curriculum leads to superior performance when combined with other strategies proposed in this section.”
    From this paper · §Approach
  • LXMERT2019 · cited 5×, 1 in Method
    “The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”
    From this paper · §Introduction
  • VL-BERT2019 · cited 5×, 1 in Method
    “The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”
    From this paper · §Introduction
  • VisualBERT2019 · cited 4×, 1 in Method
    “The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”
    From this paper · §Introduction
  • UNITER2019 · cited 4×, 1 in Method
    “Recent work has used online hard-negative mining chen2019uniter; li2019unicoder but this is costly to compute.”
    From this paper · §Approach
  • BERT2018 · cited 3×, 1 in Method
    “Internally, ViLBERT consists of two parallel BERT-style devlin2018bert models operating over image regions and text segments.”
    From this paper · §Approach
  • Unified VLP2019 · cited 3×, 1 in Method
    “The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”
    From this paper · §Introduction
  • Transformer2017 · cited 1×, 1 in Method
    “Each stream is a series of transformer blocks (TRM) vaswani2017attention connected by co-attentional transformer layers (Co-TRM) which enable information exchange between modalities.”
    From this paper · §Approach
  • COCO Captions2015 · cited 3×
    “We consider COCOcococaption and Flickr30Kplummer2015flickr30k captioning datasets for this task-group.”
    From this paper · §Vision-and-Language Tasks
  • T52019 · cited 2×
    “Advances in multi-task learning have been developed in the context of vision zhang2013robust; zhang2014facial; misra2016cross; kokkinos2017ubernet; strezoski2019many; bragman2019stochastic, language collobert2008unified;…”
    From this paper · §Related Work

Led to

  • Conceptual 12M2021 · cited 6×, 1 in Method
    “The large scale nature and the high degree of textual and visual diversity make this dataset particularly suited to V+L pre-training [55, 21, 74, 88, 48, 56].”
    From Conceptual 12M · §Vision-and-Language Pre-Training Data
  • ViTCAP2021 · cited 2×, 1 in Method
    “In total, our pre-training corpus contains 9.99.9M image-text pairs and 4.14.1M independent images, and we follow lu202012 to de-duplicate testing images exist in evaluating datasets.”
    From ViTCAP · §Experiment
Abstract

Much of vision-and-language research focuses on a small but diverse set of independent tasks and supporting datasets often studied in isolation; however, the visually-grounded language understanding skills required for success at these tasks overlap significantly. In this work, we investigate these relationships between vision-and-language tasks by developing a large-scale, multi-task training regime. Our approach culminates in a single model on 12 datasets from four broad categories of task including visual question answering, caption-based image retrieval, grounding referring expressions, and multi-modal verification. Compared to independently trained single-task models, this represents a reduction from approximately 3 billion parameters to 270 million while simultaneously improving performance by 2.05 points on average across tasks. We use our multi-task framework to perform in-depth analysis of the effect of joint training diverse tasks. Further, we show that finetuning task-specific models from our single multi-task model can lead to further improvements, achieving performance at or above the state-of-the-art.