Paper Lineage
Esc
MethodOct 2021arXiv 2110.08207cs.LG

Multitask Prompted Training Enables Zero-Shot Task Generalization

Victor Sanh, Albert Webson, Colin Raffel and 38 others

, 2020). , 2019).

From the abstract

Built on

5 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (5)

  • T52019 · cited 6×, 3 in Method
    “All models we trained are based on T5, a Transformer-based encoder-decoder language model pretrained with a masked language modeling-style objective on 1T tokens from C4 (Raffel et al. 2020).”
    From this paper · §Experimental Setup
  • True few-shot learning2021 · cited 3×, 2 in Method
    “However, this ability requires a sufficiently large model and is sensitive to the wording of its prompts (Perez et al. 2021; Zhao et al. 2021; Reynolds and McDonell 2021).”
    From this paper · §Introduction
  • GPT-32020 · cited 13×, 1 in Method
    “Finally, in explaining the success of prompts, the leading hypothesis is that models learn to understand the prompts as task instructions which help them generalize to held-out tasks (Wei et al. 2021; Mishra et al. 2021;…”
    From this paper · §Related Work
  • Prompt tuning2021 · cited 2×, 1 in Method
    “Therefore, we use Lester et al. 2021’s LM-adapted T5 model (referred to as T5+LM), produced by training T5 on 100B additional tokens from C4 on a standard language modeling objective.”
    From this paper · §Experimental Setup
Show 1 more
  • FLAN2021 · cited 8×
    “Finally, in explaining the success of prompts, the leading hypothesis is that models learn to understand the prompts as task instructions which help them generalize to held-out tasks (Wei et al. 2021; Mishra et al. 2021;…”
    From this paper · §Related Work

Led to

  • Gopher2021 · cited 1×, 1 in Method
    “Other recent LLMs include two models (FLAN and T0) fine-tuned on instructions for an array of down-stream tasks (Sanh et al. 2021; Wei et al. 2021) which improves performance to unseen tasks — these ideas are complementa…”
    From Gopher · §Background
  • InstructGPT2022 · cited 4×, 1 in Method
    “We additionally compare InstructGPT to fine-tuning 175B GPT-3 on the FLAN (Wei et al., 2021) and T0 (Sanh et al., 2021) datasets, which both consist of a variety of NLP tasks, combined with natural language instructions…”
    From InstructGPT · §Methods and experimental details
  • PaLM2022 · cited 2×
    “Any model that uses finetuning or multi-task adaptation (Wei et al. 2022a, Sanh et al. 2021) is not included in the table.”
    From PaLM · §Evaluation
  • Super-NaturalInstructions2022 · cited 5×
    “Due to the diversity of our tasks and the open-ended generation nature of our formulation,77 7 Unlike Sanh et al. 2022 and Wei et al. 2022, who evaluate their models on classification tasks via option ranking (i.e., scor…”
    From Super-NaturalInstructions · §Benchmarking Cross-Task Generalization with Sup-NatInst
  • BIG-bench2022 · cited 2×
    “BIG-bench has also been partially evaluated on other models including Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), and T0 (Sanh et al. 2022).”
    From BIG-bench · §What is in BIG-bench?
  • Emergent abilities2022 · cited 4×
    “Consider the nascent direction of enabling language models to follow natural language instructions describing a task (Wei et al. 2022a; Sanh et al. 2022; Ouyang et al. 2022, inter alia).”
    From Emergent abilities · §Discussion
  • U-PaLM2022 · cited 2×
    “A range of prior work has shown that finetuning language models on a collection of NLP tasks can improve downstream performance on a broad range of downstream tasks (Aghajanyan et al. 2021; Aribandi et al. 2022; Wei et a…”
    From U-PaLM · §Related Work
  • Flan-T5 / Flan-PaLM2022 · cited 10×
    “Further progress has been made by finetuning language models on a collection of tasks phrased as instructions, which enables models to respond better to instructions and reduces the need for few-shot exemplars (Ouyang et…”
    From Flan-T5 / Flan-PaLM · §unknown section
Abstract

Large language models have recently been shown to attain reasonable zero-shot generalization on a diverse set of tasks (Brown et al., 2020). It has been hypothesized that this is a consequence of implicit multitask learning in language models' pretraining (Radford et al., 2019). Can zero-shot generalization instead be directly induced by explicit multitask learning? To test this question at scale, we develop a system for easily mapping any natural language tasks into a human-readable prompted form. We convert a large set of supervised datasets, each with multiple prompts with diverse wording. These prompted datasets allow for benchmarking the ability of a model to perform completely held-out tasks. We fine-tune a pretrained encoder-decoder model (Raffel et al., 2020; Lester et al., 2021) on this multitask mixture covering a wide variety of tasks. The model attains strong zero-shot performance on several standard datasets, often outperforming models up to 16x its size. Further, our approach attains strong performance on a subset of tasks from the BIG-bench benchmark, outperforming models up to 6x its size. All trained models are available at https://github.com/bigscience-workshop/t-zero and all prompts are available at https://github.com/bigscience-workshop/promptsource.