Paper Lineage
Esc
MethodJun 2020arXiv 2006.06195cs.CV

Large-Scale Adversarial Training for Vision-and-Language Representation Learning

Zhe Gan, Yen-Chun Chen, Linjie Li and 3 others

We present VILLA, the first known effort on large-scale adversarial training for vision-and-language (V+L) representation learning. VILLA consists of two training stages: (i) task-agnostic adversarial pre-training; followed by (ii) task-specific adversarial finetuning.

From the abstract

Built on

8 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (8)

  • FreeLB2019 · cited 6×
    “To power efficient large-scale training, we adopt the recently proposed “free” adversarial training strategy [56, 82, 86], which obtains the gradients of parameters with almost no extra cost when computing the gradients…”
    From this paper · §Introduction
  • UNITER2019 · cited 4×
    “Inspired by the success of BERT [13] on natural language understanding, there has been a surging research interest in developing multimodal pre-training methods for vision-and-language representation learning (e.g., ViLB…”
    From this paper · §Introduction
  • “For text modality, we add adversarial perturbations to word embeddings [45, 86, 24].”
    From this paper · §Introduction
  • ViLBERT2019 · cited 3×
    “Inspired by the success of BERT [13] on natural language understanding, there has been a surging research interest in developing multimodal pre-training methods for vision-and-language representation learning (e.g., ViLB…”
    From this paper · §Introduction
Show 4 more
  • LXMERT2019 · cited 3×
    “Inspired by the success of BERT [13] on natural language understanding, there has been a surging research interest in developing multimodal pre-training methods for vision-and-language representation learning (e.g., ViLB…”
    From this paper · §Introduction
  • Visual Genome2016 · cited 2×
    “Implementation Details For UNITER experiments, we pre-train with the same four large-scale datasets used in the original model: COCO [36], Visual Genome (VG) [28], Conceptual Captions [58] and SBU Captions [49].”
    From this paper · §Experiments
  • VQA v22016 · cited 2×
    “When finetuned on downstream tasks, these pre-trained models have achieved state-of-the-art performance across diverse V+L tasks, such as Visual Question Answering (VQA) [4, 17], Visual Commonsense Reasoning (VCR) [81],…”
    From this paper · §Introduction
  • VL-BERT2019 · cited 2×
    “Compared to this two-stream architecture, recent work such as VL-BERT [60], VisualBERT [33], B2T2 [1], Unicoder-VL [30] and UNITER [12] advocate a single-stream model design, where two modalities are directly fused in ea…”
    From this paper · §Related Work

Led to

  • ALBEF2021 · cited 6×
    “The first category focuses on modelling the interactions between image and text features with transformer-based multimodal encoders [10, 11, 12, 13, 1, 14, 15, 2, 3, 16, 8, 17, 18].”
    From ALBEF · §Related Work
  • SimVLM2021 · cited 2×
    “While a variety of approaches have been proposed, a large portion of them require object detection for image region feature regression or tagging as part of the pre-training objectives (Tan & Bansal 2019; Su et al. 2020;…”
    From SimVLM · §Related Work
  • VLMo2021 · cited 3×
    “This leads to a significant speedup compared with previous models using image region features, which are extracted by an off-the-shelf object detector [30, 41, 3, 14, 25, 49].”
    From VLMo · §Experiments
Abstract

We present VILLA, the first known effort on large-scale adversarial training for vision-and-language (V+L) representation learning. VILLA consists of two training stages: (i) task-agnostic adversarial pre-training; followed by (ii) task-specific adversarial finetuning. Instead of adding adversarial perturbations on image pixels and textual tokens, we propose to perform adversarial training in the embedding space of each modality. To enable large-scale training, we adopt the "free" adversarial training strategy, and combine it with KL-divergence-based regularization to promote higher invariance in the embedding space. We apply VILLA to current best-performing V+L models, and achieve new state of the art on a wide range of tasks, including Visual Question Answering, Visual Commonsense Reasoning, Image-Text Retrieval, Referring Expression Comprehension, Visual Entailment, and NLVR2.