Paper Lineage
Esc
MethodApr 2021arXiv 2104.08691cs.CL

The Power of Scale for Parameter-Efficient Prompt Tuning

Brian Lester, Rami Al-Rfou, Noah Constant

In this work, we explore "prompt tuning", a simple yet effective mechanism for learning "soft prompts" to condition frozen language models to perform specific downstream tasks. Unlike the discrete text prompts used by GPT-3, soft prompts are learned through backpropagation and can be tuned to incorporate signal from any number of labeled examples.

From the abstract

Built on

5 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (5)

  • Prefix-Tuning2021 · cited 6×, 2 in Method
    “Li and Liang 2021 propose “prefix tuning” and show strong results on generative tasks.”
    From this paper · §Introduction
  • Adapters2019 · cited 2×, 2 in Method
    “More generally, work on task prompts is closely aligned with work on “adapters” Rebuffi et al. 2017; Houlsby et al. 2019, small bottleneck layers inserted between frozen pre-trained network layers.”
    From this paper · §Comparison to Similar Approaches
  • BART2019 · cited 1×, 1 in Method
    “Their work builds on GPT-2 Radford et al. 2019 and BART Lewis et al. 2020, while ours focuses on T5 and examines changes in performance and robustness to design choices as model size increases.”
    From this paper · §Comparison to Similar Approaches
  • T52019 · cited 6×
    “Following the “text-to-text” approach of T5 Raffel et al. 2020, we cast all tasks as text generation.”
    From this paper · §Prompt Tuning
Show 1 more
  • GPT-32020 · cited 2×
    “More recently, Brown et al. 2020 showed that prompt design (or “priming”) is surprisingly effective at modulating a frozen GPT-3 model’s behavior through text prompts.”
    From this paper · §Introduction

Led to

  • Frozen2021 · cited 2×, 1 in Method
    “Frozen is a method for grounding a large language model without changing its weights, closely related to prefix tuning [22, 19].”
    From Frozen · §The Frozen Method
  • FLAN2021 · cited 2×
    “As we’ve seen that instruction tuning improves the ability of a model to respond to instructions, it follows that, if FLAN is indeed more amenable to performing NLP tasks, then it should also achieve better performance w…”
    From FLAN · §Ablation Studies & Further Analysis
  • T02021 · cited 2×, 1 in Method
    “Therefore, we use Lester et al. 2021’s LM-adapted T5 model (referred to as T5+LM), produced by training T5 on 100B additional tokens from C4 on a standard language modeling objective.”
    From T0 · §Experimental Setup
  • Fairseq MoE LMs2021 · cited 1×, 1 in Method
    “Finally, Schick and Schütze 2021b perform few-shot learning by few-shot fine-tuning using pattern-exploiting training, whose efficiency can be improved by performing partial fine-tuning of a small number of additional ta…”
    From Fairseq MoE LMs · §Background and Related Work
Abstract

In this work, we explore "prompt tuning", a simple yet effective mechanism for learning "soft prompts" to condition frozen language models to perform specific downstream tasks. Unlike the discrete text prompts used by GPT-3, soft prompts are learned through backpropagation and can be tuned to incorporate signal from any number of labeled examples. Our end-to-end learned approach outperforms GPT-3's "few-shot" learning by a large margin. More remarkably, through ablations on model size using T5, we show that prompt tuning becomes more competitive with scale: as models exceed billions of parameters, our method "closes the gap" and matches the strong performance of model tuning (where all model weights are tuned). This finding is especially relevant in that large models are costly to share and serve, and the ability to reuse one frozen model for multiple downstream tasks can ease this burden. Our method can be seen as a simplification of the recently proposed "prefix tuning" of Li and Liang (2021), and we provide a comparison to this and other similar approaches. Finally, we show that conditioning a frozen model with soft prompts confers benefits in robustness to domain transfer, as compared to full model tuning.