Prefix-Tuning: Optimizing Continuous Prompts for Generation
Fine-tuning is the de facto way to leverage large pretrained language models to perform downstream tasks. However, it modifies all the language model parameters and therefore necessitates storing a full copy for each task.
Also cited · not yet reviewed (7)
- BART2019 · cited 3×, 1 in Method“For summarization, we compare against fine-tuning BART Lewis et al. 2020.”From this paper · §Experimental Setup
- CIDEr2014 · cited 1×, 1 in Method“We use the official evaluation script, which reports BLEU Papineni et al. 2002, NIST Belz and Reiter 2006, METEOR Lavie and Agarwal 2007, ROUGE-L Lin 2004, and CIDEr Vedantam et al. 2015.”From this paper · §Experimental Setup
- AdamW2017 · cited 1×, 1 in Method“At training time, we use the AdamW optimizer Loshchilov and Hutter 2019 and a linear learning rate scheduler, as suggested by the Hugging Face default setup.”From this paper · §Experimental Setup
- Adapters2019 · cited 3דFor example, Zhang et al. 2020a trains a “side” network that is fused with the pretrained model via summation; adapter-tuning inserts task-specific layers (adapters) between each layer of the pretrained LM Houlsby et al.…”From this paper · §Related Work
Show 3 more
- GPT-32020 · cited 3דGPT-3 Brown et al. 2020 uses manually designed prompts to adapt its generation for different tasks, and this framework is termed in-context learning.”From this paper · §Related Work
- BERT2018 · cited 2דFor extractive and abstractive summarization, researchers fine-tune masked language models (Devlin et al. 2019, e.g., BERT;) and encode-decoder models (Lewis et al. 2020, e.g., BART;) respectively Zhong et al. 2020; Liu…”From this paper · §Related Work
- T52019 · cited 2דFor table-to-text generation, Kale 2020 fine-tunes a sequence-to-sequence model (Raffel et al. 2020, T5;).”From this paper · §Related Work
Led to
- Prompt tuning2021 · cited 6×, 2 in Method“Li and Liang 2021 propose “prefix tuning” and show strong results on generative tasks.”From Prompt tuning · §Introduction
- Frozen2021 · cited 4×, 1 in Method“Frozen is a method for grounding a large language model without changing its weights, closely related to prefix tuning [22, 19].”From Frozen · §The Frozen Method
- FLAN2021 · cited 2דAs we’ve seen that instruction tuning improves the ability of a model to respond to instructions, it follows that, if FLAN is indeed more amenable to performing NLP tasks, then it should also achieve better performance w…”From FLAN · §Ablation Studies & Further Analysis
- Fairseq MoE LMs2021 · cited 1×, 1 in Method“Finally, Schick and Schütze 2021b perform few-shot learning by few-shot fine-tuning using pattern-exploiting training, whose efficiency can be improved by performing partial fine-tuning of a small number of additional ta…”From Fairseq MoE LMs · §Background and Related Work
Abstract
Fine-tuning is the de facto way to leverage large pretrained language models to perform downstream tasks. However, it modifies all the language model parameters and therefore necessitates storing a full copy for each task. In this paper, we propose prefix-tuning, a lightweight alternative to fine-tuning for natural language generation tasks, which keeps language model parameters frozen, but optimizes a small continuous task-specific vector (called the prefix). Prefix-tuning draws inspiration from prompting, allowing subsequent tokens to attend to this prefix as if it were "virtual tokens". We apply prefix-tuning to GPT-2 for table-to-text generation and to BART for summarization. We find that by learning only 0.1\% of the parameters, prefix-tuning obtains comparable performance in the full data setting, outperforms fine-tuning in low-data settings, and extrapolates better to examples with topics unseen during training.