Parameter-Efficient Transfer Learning for NLP
Fine-tuning large pre-trained models is an effective transfer mechanism in NLP. However, in the presence of many downstream tasks, fine-tuning is parameter inefficient: an entire new model is required for every task.
Also cited · not yet reviewed (2)
- BERT2018 · cited 7דTo perform classification with BERT, we follow the approach in Devlin et al. 2018.”From this paper · §Experiments
- Transformer2017 · cited 3דRecent state-of-the-art results on question answering (Rajpurkar et al. 2016) and text classification (Wang et al. 2018) have been attained by fine-tuning a Transformer network (Vaswani et al. 2017) with a Masked Languag…”From this paper · §Related Work
Led to
- T52019 · cited 3דThis synergy has resulted in a great deal of recent work developing transfer learning methodology for NLP, which has produced a wide landscape of pre-training objectives (Howard and Ruder 2018; Devlin et al. 2018; Yang e…”From T5 · §Introduction
- Prefix-Tuning2021 · cited 3דFor example, Zhang et al. 2020a trains a “side” network that is fused with the pretrained model via summation; adapter-tuning inserts task-specific layers (adapters) between each layer of the pretrained LM Houlsby et al.…”From Prefix-Tuning · §Related Work
- Prompt tuning2021 · cited 2×, 2 in Method“More generally, work on task prompts is closely aligned with work on “adapters” Rebuffi et al. 2017; Houlsby et al. 2019, small bottleneck layers inserted between frozen pre-trained network layers.”From Prompt tuning · §Comparison to Similar Approaches
- Fairseq MoE LMs2021 · cited 1×, 1 in Method“Finally, Schick and Schütze 2021b perform few-shot learning by few-shot fine-tuning using pattern-exploiting training, whose efficiency can be improved by performing partial fine-tuning of a small number of additional ta…”From Fairseq MoE LMs · §Background and Related Work
Abstract
Fine-tuning large pre-trained models is an effective transfer mechanism in NLP. However, in the presence of many downstream tasks, fine-tuning is parameter inefficient: an entire new model is required for every task. As an alternative, we propose transfer with adapter modules. Adapter modules yield a compact and extensible model; they add only a few trainable parameters per task, and new tasks can be added without revisiting previous ones. The parameters of the original network remain fixed, yielding a high degree of parameter sharing. To demonstrate adapter's effectiveness, we transfer the recently proposed BERT Transformer model to 26 diverse text classification tasks, including the GLUE benchmark. Adapters attain near state-of-the-art performance, whilst adding only a few parameters per task. On GLUE, we attain within 0.4% of the performance of full fine-tuning, adding only 3.6% parameters per task. By contrast, fine-tuning trains 100% of the parameters per task.