Paper Lineage
Esc
MethodSep 2019arXiv 1909.11764cs.CL

FreeLB: Enhanced Adversarial Training for Natural Language Understanding

Chen Zhu, Yu Cheng, Zhe Gan and 3 others

Adversarial training, which minimizes the maximal risk for label-preserving input perturbations, has proved to be effective for improving the generalization of language models. In this work, we propose a novel adversarial training algorithm, FreeLB, that promotes higher invariance in the embedding space, by adding adversarial perturbations to word embeddings and minimizing the resultant adversarial risk inside different regions around input samples.

From the abstract

Built on

3 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (3)

  • RoBERTa2019 · cited 4×
    “CommonsenseQA Similar to the training strategy in Liu et al. 2019b, we construct five inputs for each question by concatenating the question and each answer separately, then encode each input with the representation of t…”
    From this paper · §Experiments
  • BERT2018 · cited 2×
    “For BERT-base, we use the HuggingFace implementation44 4 https://github.com/huggingface/pytorch-transformers, and follow the single-task finetuning procedure as in Devlin et al. 2019.”
    From this paper · §Experiments
  • ALBERT2019 · cited 2×
    “To further explore its ability to improve more sophisticated language models, we apply FreeLB to the fine-tuning stage of ALBERT-xxlarge-v2 (Lan et al. 2020) model on the dev set of GLUE.”
    From this paper · §Experiments

Led to

  • VILLA2020 · cited 6×
    “To power efficient large-scale training, we adopt the recently proposed “free” adversarial training strategy [56, 82, 86], which obtains the gradients of parameters with almost no extra cost when computing the gradients…”
    From VILLA · §Introduction
Abstract

Adversarial training, which minimizes the maximal risk for label-preserving input perturbations, has proved to be effective for improving the generalization of language models. In this work, we propose a novel adversarial training algorithm, FreeLB, that promotes higher invariance in the embedding space, by adding adversarial perturbations to word embeddings and minimizing the resultant adversarial risk inside different regions around input samples. To validate the effectiveness of the proposed approach, we apply it to Transformer-based models for natural language understanding and commonsense reasoning tasks. Experiments on the GLUE benchmark show that when applied only to the finetuning stage, it is able to improve the overall test scores of BERT-base model from 78.3 to 79.4, and RoBERTa-large model from 88.5 to 88.8. In addition, the proposed approach achieves state-of-the-art single-model test accuracies of 85.44\% and 67.75\% on ARC-Easy and ARC-Challenge. Experiments on CommonsenseQA benchmark further demonstrate that FreeLB can be generalized and boost the performance of RoBERTa-large model on other tasks as well. Code is available at \url{https://github.com/zhuchen03/FreeLB .