Large-Scale Adversarial Training for Vision-and-Language Representation Learning
We present VILLA, the first known effort on large-scale adversarial training for vision-and-language (V+L) representation learning. VILLA consists of two training stages: (i) task-agnostic adversarial pre-training; followed by (ii) task-specific adversarial finetuning.
Also cited · not yet reviewed (8)
- FreeLB2019 · cited 6דTo power efficient large-scale training, we adopt the recently proposed “free” adversarial training strategy [56, 82, 86], which obtains the gradients of parameters with almost no extra cost when computing the gradients…”From this paper · §Introduction
- UNITER2019 · cited 4דInspired by the success of BERT [13] on natural language understanding, there has been a surging research interest in developing multimodal pre-training methods for vision-and-language representation learning (e.g., ViLB…”From this paper · §Introduction
- Bottom-Up Top-Down attention2017 · cited 3דFor text modality, we add adversarial perturbations to word embeddings [45, 86, 24].”From this paper · §Introduction
- ViLBERT2019 · cited 3דInspired by the success of BERT [13] on natural language understanding, there has been a surging research interest in developing multimodal pre-training methods for vision-and-language representation learning (e.g., ViLB…”From this paper · §Introduction
Show 4 more
- LXMERT2019 · cited 3דInspired by the success of BERT [13] on natural language understanding, there has been a surging research interest in developing multimodal pre-training methods for vision-and-language representation learning (e.g., ViLB…”From this paper · §Introduction
- Visual Genome2016 · cited 2דImplementation Details For UNITER experiments, we pre-train with the same four large-scale datasets used in the original model: COCO [36], Visual Genome (VG) [28], Conceptual Captions [58] and SBU Captions [49].”From this paper · §Experiments
- VQA v22016 · cited 2דWhen finetuned on downstream tasks, these pre-trained models have achieved state-of-the-art performance across diverse V+L tasks, such as Visual Question Answering (VQA) [4, 17], Visual Commonsense Reasoning (VCR) [81],…”From this paper · §Introduction
- VL-BERT2019 · cited 2דCompared to this two-stream architecture, recent work such as VL-BERT [60], VisualBERT [33], B2T2 [1], Unicoder-VL [30] and UNITER [12] advocate a single-stream model design, where two modalities are directly fused in ea…”From this paper · §Related Work
Led to
- ALBEF2021 · cited 6דThe first category focuses on modelling the interactions between image and text features with transformer-based multimodal encoders [10, 11, 12, 13, 1, 14, 15, 2, 3, 16, 8, 17, 18].”From ALBEF · §Related Work
- SimVLM2021 · cited 2דWhile a variety of approaches have been proposed, a large portion of them require object detection for image region feature regression or tagging as part of the pre-training objectives (Tan & Bansal 2019; Su et al. 2020;…”From SimVLM · §Related Work
- VLMo2021 · cited 3דThis leads to a significant speedup compared with previous models using image region features, which are extracted by an off-the-shelf object detector [30, 41, 3, 14, 25, 49].”From VLMo · §Experiments
Abstract
We present VILLA, the first known effort on large-scale adversarial training for vision-and-language (V+L) representation learning. VILLA consists of two training stages: (i) task-agnostic adversarial pre-training; followed by (ii) task-specific adversarial finetuning. Instead of adding adversarial perturbations on image pixels and textual tokens, we propose to perform adversarial training in the embedding space of each modality. To enable large-scale training, we adopt the "free" adversarial training strategy, and combine it with KL-divergence-based regularization to promote higher invariance in the embedding space. We apply VILLA to current best-performing V+L models, and achieve new state of the art on a wide range of tasks, including Visual Question Answering, Visual Commonsense Reasoning, Image-Text Retrieval, Referring Expression Comprehension, Visual Entailment, and NLVR2.