Paper Lineage
Esc
AnalysisNov 2020arXiv 2011.15124cs.CL

Multimodal Pretraining Unmasked: A Meta-Analysis and a Unified Framework of Vision-and-Language BERTs

Emanuele Bugliarello, Ryan Cotterell, Naoaki Okazaki, Desmond Elliott

Large-scale pretraining and task-specific fine-tuning is now the standard methodology for many tasks in computer vision and natural language processing. Recently, a multitude of methods have been proposed for pretraining vision and language BERTs to tackle challenges at the intersection of these two key areas of AI.

From the abstract

Built on

19 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (19)

  • MS COCO2014 · cited 2×, 2 in Method
    “COCO (Lin et al. 2014) or VQA (Antol et al. 2015), where the images are strongly-associated with crowdsourced captions or question--answer pairs.”
    From this paper · §Experimental Setup
  • Transformer2017 · cited 2×, 2 in Method
    “This is similar to the input masking applied in autoregressive Transformer decoders (Vaswani et al. 2017).”
    From this paper · §A Unified Framework
  • LXMERT2019 · cited 3×, 1 in Method
    “Each model is initialized with the parameters of BERT, following the approaches described in the original papers.88 8 Only Tan and Bansal 2019 reported slightly better performance when pretraining from scratch but they r…”
    From this paper · §Experimental Setup
  • VQA v22016 · cited 2×, 1 in Method
    “We consider the most common tasks used to evaluate V&L BERTs, spanning four groups: vocab-based VQA (Goyal et al. 2017; Hudson and Manning 2019), image–text retrieval (Lin et al. 2014; Plummer et al. 2015), referring exp…”
    From this paper · §Experimental Setup
Show 15 more
  • VQA2015 · cited 1×, 1 in Method
    “COCO (Lin et al. 2014) or VQA (Antol et al. 2015), where the images are strongly-associated with crowdsourced captions or question--answer pairs.”
    From this paper · §Experimental Setup
  • Flickr30k Entities2015 · cited 1×, 1 in Method
    “We consider the most common tasks used to evaluate V&L BERTs, spanning four groups: vocab-based VQA (Goyal et al. 2017; Hudson and Manning 2019), image–text retrieval (Lin et al. 2014; Plummer et al. 2015), referring exp…”
    From this paper · §Experimental Setup
  • Faster R-CNN2015 · cited 1×, 1 in Method
    “V&L BERTs typically extract image features using a Faster R-CNN (Ren et al. 2015) trained on the Visual Genome dataset (VG; Krishna et al. 2017), either with a ResNet-101101 (He et al. 2016) or a ResNeXT-152152 backbone…”
    From this paper · §Experimental Setup
  • ResNet2015 · cited 1×, 1 in Method
    “V&L BERTs typically extract image features using a Faster R-CNN (Ren et al. 2015) trained on the Visual Genome dataset (VG; Krishna et al. 2017), either with a ResNet-101101 (He et al. 2016) or a ResNeXT-152152 backbone…”
    From this paper · §Experimental Setup
  • Visual Genome2016 · cited 1×, 1 in Method
    “V&L BERTs typically extract image features using a Faster R-CNN (Ren et al. 2015) trained on the Visual Genome dataset (VG; Krishna et al. 2017), either with a ResNet-101101 (He et al. 2016) or a ResNeXT-152152 backbone…”
    From this paper · §Experimental Setup
  • ResNeXt2016 · cited 1×, 1 in Method
    “V&L BERTs typically extract image features using a Faster R-CNN (Ren et al. 2015) trained on the Visual Genome dataset (VG; Krishna et al. 2017), either with a ResNet-101101 (He et al. 2016) or a ResNeXT-152152 backbone…”
    From this paper · §Experimental Setup
  • Bottom-Up Top-Down attention2017 · cited 1×, 1 in Method
    “Our models are trained with 3636 regions of interest extracted by a Faster R-CNN with a ResNet-101101 backbone (Anderson et al. 2018).”
    From this paper · §Experimental Setup
  • NLVR22018 · cited 1×, 1 in Method
    “We consider the most common tasks used to evaluate V&L BERTs, spanning four groups: vocab-based VQA (Goyal et al. 2017; Hudson and Manning 2019), image–text retrieval (Lin et al. 2014; Plummer et al. 2015), referring exp…”
    From this paper · §Experimental Setup
  • UNITER2019 · cited 7×
    “The majority of V&L BERTs follow the single-stream paradigm (Su et al. 2020; Li et al. 2019; Chen et al. 2020; Li et al. 2020a; Zhou et al. 2020; Lin et al. 2020; Li et al. 2020b).”
    From this paper · §Vision-and-Language BERTs
  • ViLBERT2019 · cited 5×
    “ViLBERT (Lu et al. 2019), LXMERT (Tan and Bansal 2019), and ERNIE-ViL (Yu et al. 2021)33 3 ERNIE-ViL uses the dual-stream ViLBERT encoder. are based on a dual-stream paradigm.”
    From this paper · §Vision-and-Language BERTs
  • BERT2018 · cited 3×
    “In pursuit of this goal, many pretrained V&L models have been proposed in the last year, inspired by the success of pretraining in both computer vision (Sharif Razavian et al. 2014) and natural language processing (Devli…”
    From this paper · §Introduction
  • VL-BERT2019 · cited 3×
    “While most V&L BERTs follow this paradigm, some studies find beneficial to jointly learn the visual encoder with language Su et al. 2020; Huang et al. 2020; Radford et al. 2021; Kim et al. 2021.”
    From this paper · §Results
  • VisualBERT2019 · cited 2×
    “While most approaches present very similar ways to embed spatial locations, VL-BERT relies on a more complex geometry embedding and they are, instead, missing in VisualBERT Li et al. 2019.”
    From this paper · §Vision-and-Language BERTs
  • Unicoder-VL2019 · cited 2×
    “The majority of V&L BERTs follow the single-stream paradigm (Su et al. 2020; Li et al. 2019; Chen et al. 2020; Li et al. 2020a; Zhou et al. 2020; Lin et al. 2020; Li et al. 2020b).”
    From this paper · §Vision-and-Language BERTs
  • Unified VLP2019 · cited 2×
    “The majority of V&L BERTs follow the single-stream paradigm (Su et al. 2020; Li et al. 2019; Chen et al. 2020; Li et al. 2020a; Zhou et al. 2020; Lin et al. 2020; Li et al. 2020b).”
    From this paper · §Vision-and-Language BERTs

Led to

  • ViLT2021 · cited 2×, 2 in Method
    “Bugliarello et al. 2020 classifies interaction schema into two categories: (1) single-stream approaches (e.g., VisualBERT (Li et al. 2019), UNITER (Chen et al. 2019)) where layers collectively operate on a concatenation…”
    From ViLT · §Background
Abstract

Large-scale pretraining and task-specific fine-tuning is now the standard methodology for many tasks in computer vision and natural language processing. Recently, a multitude of methods have been proposed for pretraining vision and language BERTs to tackle challenges at the intersection of these two key areas of AI. These models can be categorised into either single-stream or dual-stream encoders. We study the differences between these two categories, and show how they can be unified under a single theoretical framework. We then conduct controlled experiments to discern the empirical differences between five V&L BERTs. Our experiments show that training data and hyperparameters are responsible for most of the differences between the reported results, but they also reveal that the embedding layer plays a crucial role in these massive models.