12-in-1: Multi-Task Vision and Language Representation Learning
Much of vision-and-language research focuses on a small but diverse set of independent tasks and supporting datasets often studied in isolation; however, the visually-grounded language understanding skills required for success at these tasks overlap significantly. In this work, we investigate these relationships between vision-and-language tasks by developing a large-scale, multi-task training regime.
Also cited · not yet reviewed (14)
- ViLBERT2019 · cited 9×, 4 in Method“The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”From this paper · §Introduction
- Unicoder-VL2019 · cited 6×, 2 in Method“The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”From this paper · §Introduction
- decaNLP2018 · cited 5×, 2 in Method“Inspired by prior multi-task literature bengio2009curriculum mccann2018natural, we experimented with both curriculum and anti-curriculum strategies based on task difficulty.”From this paper · §Approach
- Bottom-Up Top-Down attention2017 · cited 2×, 2 in Method“At the interface level, ViLBERT takes as input an image II and text segment QQ represented as the sequence {\{IMG,v1,…,v𝒯,,v_{1},\dots,v_{T}, CLS, w1,…,wT,w_{1},\dots,w_{T}, SEP}\} where {vi}i=1𝒯\{v_{i}\}_{i=1}^{T} are…”From this paper · §Approach
Show 10 more
- Multi-task hierarchical VL2018 · cited 6×, 1 in Method“Different from previous observation mccann2018natural; nguyen2019multi, we found that using no curriculum leads to superior performance when combined with other strategies proposed in this section.”From this paper · §Approach
- LXMERT2019 · cited 5×, 1 in Method“The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”From this paper · §Introduction
- VL-BERT2019 · cited 5×, 1 in Method“The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”From this paper · §Introduction
- VisualBERT2019 · cited 4×, 1 in Method“The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”From this paper · §Introduction
- UNITER2019 · cited 4×, 1 in Method“Recent work has used online hard-negative mining chen2019uniter; li2019unicoder but this is costly to compute.”From this paper · §Approach
- BERT2018 · cited 3×, 1 in Method“Internally, ViLBERT consists of two parallel BERT-style devlin2018bert models operating over image regions and text segments.”From this paper · §Approach
- Unified VLP2019 · cited 3×, 1 in Method“The recent rise of general architectures for vision-and-language lu2019vilbert; tan2019lxmert; li2019visualbert; alberti2019fusion; li2019unicoder; su2019vl; zhou2019unified reduces the architectural differences across t…”From this paper · §Introduction
- Transformer2017 · cited 1×, 1 in Method“Each stream is a series of transformer blocks (TRM) vaswani2017attention connected by co-attentional transformer layers (Co-TRM) which enable information exchange between modalities.”From this paper · §Approach
- COCO Captions2015 · cited 3דWe consider COCOcococaption and Flickr30Kplummer2015flickr30k captioning datasets for this task-group.”From this paper · §Vision-and-Language Tasks
- T52019 · cited 2דAdvances in multi-task learning have been developed in the context of vision zhang2013robust; zhang2014facial; misra2016cross; kokkinos2017ubernet; strezoski2019many; bragman2019stochastic, language collobert2008unified;…”From this paper · §Related Work
Led to
- Conceptual 12M2021 · cited 6×, 1 in Method“The large scale nature and the high degree of textual and visual diversity make this dataset particularly suited to V+L pre-training [55, 21, 74, 88, 48, 56].”From Conceptual 12M · §Vision-and-Language Pre-Training Data
- ViTCAP2021 · cited 2×, 1 in Method“In total, our pre-training corpus contains 9.99.9M image-text pairs and 4.14.1M independent images, and we follow lu202012 to de-duplicate testing images exist in evaluating datasets.”From ViTCAP · §Experiment
Abstract
Much of vision-and-language research focuses on a small but diverse set of independent tasks and supporting datasets often studied in isolation; however, the visually-grounded language understanding skills required for success at these tasks overlap significantly. In this work, we investigate these relationships between vision-and-language tasks by developing a large-scale, multi-task training regime. Our approach culminates in a single model on 12 datasets from four broad categories of task including visual question answering, caption-based image retrieval, grounding referring expressions, and multi-modal verification. Compared to independently trained single-task models, this represents a reduction from approximately 3 billion parameters to 270 million while simultaneously improving performance by 2.05 points on average across tasks. We use our multi-task framework to perform in-depth analysis of the effect of joint training diverse tasks. Further, we show that finetuning task-specific models from our single multi-task model can lead to further improvements, achieving performance at or above the state-of-the-art.