Paper Lineage
Esc
DatasetApr 2022arXiv 2204.07705cs.CL

Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks

Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi and 37 others

How well can NLP models generalize to a variety of unseen tasks when provided with task instructions? To address this question, we first introduce Super-NaturalInstructions, a benchmark of 1,616 diverse NLP tasks and their expert-written instructions.

From the abstract

Built on

6 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (6)

  • Natural Instructions2021 · cited 6×
    “Recent literature has been motivated by building models that are generalizable across a variety of NLP tasks, when prompted with either a few examples Ye and Ren 2021; Bragg et al. 2021 or language definitions Efrat and…”
    From this paper · §Related Work
  • T02021 · cited 5×
    “Due to the diversity of our tasks and the open-ended generation nature of our formulation,77 7 Unlike Sanh et al. 2022 and Wei et al. 2022, who evaluate their models on classification tasks via option ranking (i.e., scor…”
    From this paper · §Benchmarking Cross-Task Generalization with Sup-NatInst
  • T52019 · cited 4×
    “We conduct our experiments and analysis based on the T5 model Raffel et al. 2020.”
    From this paper · §Tkk-Instruct: Learning to Follow Instructions at Scale
  • FLAN2021 · cited 3×
    “Due to the diversity of our tasks and the open-ended generation nature of our formulation,77 7 Unlike Sanh et al. 2022 and Wei et al. 2022, who evaluate their models on classification tasks via option ranking (i.e., scor…”
    From this paper · §Benchmarking Cross-Task Generalization with Sup-NatInst
Show 2 more
  • InstructGPT2022 · cited 3×
    “Finally, the well-adopted InstructGPT model Ouyang et al. 2022 is partially enabled by a large dataset of prompts that are collected via various synthetic data augmentation which, unfortunately, is not publicly available…”
    From this paper · §Related Work
  • GPT-32020 · cited 2×
    “Additionally, we evaluate GPT-3 Brown et al. 2020, a 175175B-parameter autoregressive LM that has shown remarkable ability in following demonstrations provided in its prompt.”
    From this paper · §Benchmarking Cross-Task Generalization with Sup-NatInst

Led to

  • Flan-T5 / Flan-PaLM2022 · cited 7×
    “For instance, Flan-PaLM’s improved reasoning abilities enable it to leverage CoT and self-consistency (Wang et al. 2022c) to achieve 75.2% on Massive Multi-task Language Understanding (Hendrycks et al. 2020, MMLU;).”
    From Flan-T5 / Flan-PaLM · §unknown section
Abstract

How well can NLP models generalize to a variety of unseen tasks when provided with task instructions? To address this question, we first introduce Super-NaturalInstructions, a benchmark of 1,616 diverse NLP tasks and their expert-written instructions. Our collection covers 76 distinct task types, including but not limited to classification, extraction, infilling, sequence tagging, text rewriting, and text composition. This large and diverse collection of tasks enables rigorous benchmarking of cross-task generalization under instructions -- training models to follow instructions on a subset of tasks and evaluating them on the remaining unseen ones. Furthermore, we build Tk-Instruct, a transformer model trained to follow a variety of in-context instructions (plain language task definitions or k-shot examples). Our experiments show that Tk-Instruct outperforms existing instruction-following models such as InstructGPT by over 9% on our benchmark despite being an order of magnitude smaller. We further analyze generalization as a function of various scaling parameters, such as the number of observed tasks, the number of instances per task, and model sizes. We hope our dataset and model facilitate future progress towards more general-purpose NLP models.