Finetuned Language Models Are Zero-Shot Learners
This paper explores a simple method for improving the zero-shot learning abilities of language models. We show that instruction tuning -- finetuning language models on a collection of tasks described via instructions -- substantially improves zero-shot performance on unseen tasks.
Also cited · not yet reviewed (4)
- GPT-32020 · cited 8דLanguage models (LMs) at scale, such as GPT-3 (Brown et al. 2020), have been shown to perform few-shot learning remarkably well.”From this paper · §Introduction
- T52019 · cited 3דTo balance the different sizes of datasets, we limit the number of training examples per dataset to 30k and follow the examples-proportional mixing scheme (Raffel et al. 2020) with a mixing rate maximum of 3k.22 2 In thi…”From this paper · §FLAN: Instruction Tuning Improves Zero-Shot Learning
- Prefix-Tuning2021 · cited 2דAs we’ve seen that instruction tuning improves the ability of a model to respond to instructions, it follows that, if FLAN is indeed more amenable to performing NLP tasks, then it should also achieve better performance w…”From this paper · §Ablation Studies & Further Analysis
- Prompt tuning2021 · cited 2דAs we’ve seen that instruction tuning improves the ability of a model to respond to instructions, it follows that, if FLAN is indeed more amenable to performing NLP tasks, then it should also achieve better performance w…”From this paper · §Ablation Studies & Further Analysis
Led to
- T02021 · cited 8דFinally, in explaining the success of prompts, the leading hypothesis is that models learn to understand the prompts as task instructions which help them generalize to held-out tasks (Wei et al. 2021; Mishra et al. 2021;…”From T0 · §Related Work
- Gopher2021 · cited 1×, 1 in Method“Other recent LLMs include two models (FLAN and T0) fine-tuned on instructions for an array of down-stream tasks (Sanh et al. 2021; Wei et al. 2021) which improves performance to unseen tasks — these ideas are complementa…”From Gopher · §Background
- GLaM2021 · cited 2דGPT-3 (Brown et al. 2020) and related work (Shoeybi et al. 2019; Lieber et al. 2021; Wei et al. 2021) demonstrated that scaling up language models greatly improves task-agnostic, few-shot performance.”From GLaM · §Related Work
- Megatron-Turing NLG2022 · cited 2דBoth T0 [59] and FLAN [70] have taken this path and have shown that such an approach can improve zero-shot learning capabilities of language models.”From Megatron-Turing NLG · §Related Works
- InstructGPT2022 · cited 4×, 1 in Method“We additionally compare InstructGPT to fine-tuning 175B GPT-3 on the FLAN (Wei et al., 2021) and T0 (Sanh et al., 2021) datasets, which both consist of a variety of NLP tasks, combined with natural language instructions…”From InstructGPT · §Methods and experimental details
- PaLM2022 · cited 5דAny model that uses finetuning or multi-task adaptation (Wei et al. 2022a, Sanh et al. 2021) is not included in the table.”From PaLM · §Evaluation
- Super-NaturalInstructions2022 · cited 3דDue to the diversity of our tasks and the open-ended generation nature of our formulation,77 7 Unlike Sanh et al. 2022 and Wei et al. 2022, who evaluate their models on classification tasks via option ranking (i.e., scor…”From Super-NaturalInstructions · §Benchmarking Cross-Task Generalization with Sup-NatInst
- OPT2022 · cited 2דRecent efforts have shown gains by fine-tuning models to directly respond to instruction-style prompting Wei et al. 2021; Min et al. 2021; Sanh et al. 2021; Ouyang et al. 2022.”From OPT · §Related Work
- Emergent abilities2022 · cited 4דConsider the nascent direction of enabling language models to follow natural language instructions describing a task (Wei et al. 2022a; Sanh et al. 2022; Ouyang et al. 2022, inter alia).”From Emergent abilities · §Discussion
- U-PaLM2022 · cited 2דA range of prior work has shown that finetuning language models on a collection of NLP tasks can improve downstream performance on a broad range of downstream tasks (Aghajanyan et al. 2021; Aribandi et al. 2022; Wei et a…”From U-PaLM · §Related Work
- Flan-T5 / Flan-PaLM2022 · cited 15דWe call this finetuning procedure Flan (Wei et al. 2021, Finetuning language models;) and prepend ‘‘Flan’’ to the resulting finetuned models (e.g., Flan-PaLM).22 2 We use “Flan” to refer to our finetuning procedure. “FLA…”From Flan-T5 / Flan-PaLM · §unknown section
Abstract
This paper explores a simple method for improving the zero-shot learning abilities of language models. We show that instruction tuning -- finetuning language models on a collection of tasks described via instructions -- substantially improves zero-shot performance on unseen tasks. We take a 137B parameter pretrained language model and instruction-tune it on over 60 NLP tasks verbalized via natural language instruction templates. We evaluate this instruction-tuned model, which we call FLAN, on unseen task types. FLAN substantially improves the performance of its unmodified counterpart and surpasses zero-shot 175B GPT-3 on 20 of 25 tasks that we evaluate. FLAN even outperforms few-shot GPT-3 by a large margin on ANLI, RTE, BoolQ, AI2-ARC, OpenbookQA, and StoryCloze. Ablation studies reveal that number of finetuning datasets, model scale, and natural language instructions are key to the success of instruction tuning.