Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks
How well can NLP models generalize to a variety of unseen tasks when provided with task instructions? To address this question, we first introduce Super-NaturalInstructions, a benchmark of 1,616 diverse NLP tasks and their expert-written instructions.
Also cited · not yet reviewed (6)
- Natural Instructions2021 · cited 6דRecent literature has been motivated by building models that are generalizable across a variety of NLP tasks, when prompted with either a few examples Ye and Ren 2021; Bragg et al. 2021 or language definitions Efrat and…”From this paper · §Related Work
- T02021 · cited 5דDue to the diversity of our tasks and the open-ended generation nature of our formulation,77 7 Unlike Sanh et al. 2022 and Wei et al. 2022, who evaluate their models on classification tasks via option ranking (i.e., scor…”From this paper · §Benchmarking Cross-Task Generalization with Sup-NatInst
- T52019 · cited 4דWe conduct our experiments and analysis based on the T5 model Raffel et al. 2020.”From this paper · §Tkk-Instruct: Learning to Follow Instructions at Scale
- FLAN2021 · cited 3דDue to the diversity of our tasks and the open-ended generation nature of our formulation,77 7 Unlike Sanh et al. 2022 and Wei et al. 2022, who evaluate their models on classification tasks via option ranking (i.e., scor…”From this paper · §Benchmarking Cross-Task Generalization with Sup-NatInst
Show 2 more
- InstructGPT2022 · cited 3דFinally, the well-adopted InstructGPT model Ouyang et al. 2022 is partially enabled by a large dataset of prompts that are collected via various synthetic data augmentation which, unfortunately, is not publicly available…”From this paper · §Related Work
- GPT-32020 · cited 2דAdditionally, we evaluate GPT-3 Brown et al. 2020, a 175175B-parameter autoregressive LM that has shown remarkable ability in following demonstrations provided in its prompt.”From this paper · §Benchmarking Cross-Task Generalization with Sup-NatInst
Led to
- Flan-T5 / Flan-PaLM2022 · cited 7דFor instance, Flan-PaLM’s improved reasoning abilities enable it to leverage CoT and self-consistency (Wang et al. 2022c) to achieve 75.2% on Massive Multi-task Language Understanding (Hendrycks et al. 2020, MMLU;).”From Flan-T5 / Flan-PaLM · §unknown section
Abstract
How well can NLP models generalize to a variety of unseen tasks when provided with task instructions? To address this question, we first introduce Super-NaturalInstructions, a benchmark of 1,616 diverse NLP tasks and their expert-written instructions. Our collection covers 76 distinct task types, including but not limited to classification, extraction, infilling, sequence tagging, text rewriting, and text composition. This large and diverse collection of tasks enables rigorous benchmarking of cross-task generalization under instructions -- training models to follow instructions on a subset of tasks and evaluating them on the remaining unseen ones. Furthermore, we build Tk-Instruct, a transformer model trained to follow a variety of in-context instructions (plain language task definitions or k-shot examples). Our experiments show that Tk-Instruct outperforms existing instruction-following models such as InstructGPT by over 9% on our benchmark despite being an order of magnitude smaller. We further analyze generalization as a function of various scaling parameters, such as the number of observed tasks, the number of instances per task, and model sizes. We hope our dataset and model facilitate future progress towards more general-purpose NLP models.