Multitask Prompted Training Enables Zero-Shot Task Generalization
, 2020). , 2019).
Also cited · not yet reviewed (5)
- T52019 · cited 6×, 3 in Method“All models we trained are based on T5, a Transformer-based encoder-decoder language model pretrained with a masked language modeling-style objective on 1T tokens from C4 (Raffel et al. 2020).”From this paper · §Experimental Setup
- True few-shot learning2021 · cited 3×, 2 in Method“However, this ability requires a sufficiently large model and is sensitive to the wording of its prompts (Perez et al. 2021; Zhao et al. 2021; Reynolds and McDonell 2021).”From this paper · §Introduction
- GPT-32020 · cited 13×, 1 in Method“Finally, in explaining the success of prompts, the leading hypothesis is that models learn to understand the prompts as task instructions which help them generalize to held-out tasks (Wei et al. 2021; Mishra et al. 2021;…”From this paper · §Related Work
- Prompt tuning2021 · cited 2×, 1 in Method“Therefore, we use Lester et al. 2021’s LM-adapted T5 model (referred to as T5+LM), produced by training T5 on 100B additional tokens from C4 on a standard language modeling objective.”From this paper · §Experimental Setup
Show 1 more
- FLAN2021 · cited 8דFinally, in explaining the success of prompts, the leading hypothesis is that models learn to understand the prompts as task instructions which help them generalize to held-out tasks (Wei et al. 2021; Mishra et al. 2021;…”From this paper · §Related Work
Led to
- Gopher2021 · cited 1×, 1 in Method“Other recent LLMs include two models (FLAN and T0) fine-tuned on instructions for an array of down-stream tasks (Sanh et al. 2021; Wei et al. 2021) which improves performance to unseen tasks — these ideas are complementa…”From Gopher · §Background
- InstructGPT2022 · cited 4×, 1 in Method“We additionally compare InstructGPT to fine-tuning 175B GPT-3 on the FLAN (Wei et al., 2021) and T0 (Sanh et al., 2021) datasets, which both consist of a variety of NLP tasks, combined with natural language instructions…”From InstructGPT · §Methods and experimental details
- PaLM2022 · cited 2דAny model that uses finetuning or multi-task adaptation (Wei et al. 2022a, Sanh et al. 2021) is not included in the table.”From PaLM · §Evaluation
- Super-NaturalInstructions2022 · cited 5דDue to the diversity of our tasks and the open-ended generation nature of our formulation,77 7 Unlike Sanh et al. 2022 and Wei et al. 2022, who evaluate their models on classification tasks via option ranking (i.e., scor…”From Super-NaturalInstructions · §Benchmarking Cross-Task Generalization with Sup-NatInst
- BIG-bench2022 · cited 2דBIG-bench has also been partially evaluated on other models including Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), and T0 (Sanh et al. 2022).”From BIG-bench · §What is in BIG-bench?
- Emergent abilities2022 · cited 4דConsider the nascent direction of enabling language models to follow natural language instructions describing a task (Wei et al. 2022a; Sanh et al. 2022; Ouyang et al. 2022, inter alia).”From Emergent abilities · §Discussion
- U-PaLM2022 · cited 2דA range of prior work has shown that finetuning language models on a collection of NLP tasks can improve downstream performance on a broad range of downstream tasks (Aghajanyan et al. 2021; Aribandi et al. 2022; Wei et a…”From U-PaLM · §Related Work
- Flan-T5 / Flan-PaLM2022 · cited 10דFurther progress has been made by finetuning language models on a collection of tasks phrased as instructions, which enables models to respond better to instructions and reduces the need for few-shot exemplars (Ouyang et…”From Flan-T5 / Flan-PaLM · §unknown section
Abstract
Large language models have recently been shown to attain reasonable zero-shot generalization on a diverse set of tasks (Brown et al., 2020). It has been hypothesized that this is a consequence of implicit multitask learning in language models' pretraining (Radford et al., 2019). Can zero-shot generalization instead be directly induced by explicit multitask learning? To test this question at scale, we develop a system for easily mapping any natural language tasks into a human-readable prompted form. We convert a large set of supervised datasets, each with multiple prompts with diverse wording. These prompted datasets allow for benchmarking the ability of a model to perform completely held-out tasks. We fine-tune a pretrained encoder-decoder model (Raffel et al., 2020; Lester et al., 2021) on this multitask mixture covering a wide variety of tasks. The model attains strong zero-shot performance on several standard datasets, often outperforming models up to 16x its size. Further, our approach attains strong performance on a subset of tasks from the BIG-bench benchmark, outperforming models up to 6x its size. All trained models are available at https://github.com/bigscience-workshop/t-zero and all prompts are available at https://github.com/bigscience-workshop/promptsource.