Paper Lineage
Esc
AnalysisJun 2022arXiv 2206.07682cs.CL

Emergent Abilities of Large Language Models

Jason Wei, Yi Tay, Rishi Bommasani and 13 others

Scaling up language models has been shown to predictably improve performance and sample efficiency on a wide range of downstream tasks. This paper instead discusses an unpredictable phenomenon that we refer to as emergent abilities of large language models.

From the abstract

Built on

12 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (12)

  • GPT-32020 · cited 9×
    “It is now well-known that increasing the scale of language models (e.g., training compute, model parameters, etc.) can lead to better performance and sample efficiency on a range of downstream NLP tasks (Devlin et al. 20…”
    From this paper · §Introduction
  • Chinchilla2022 · cited 6×
    “In many cases, the effect of scale on performance can often be methodologically predicted via scaling laws—for example, scaling curves for cross-entropy loss have been shown to empirically span more than seven orders of…”
    From this paper · §Introduction
  • Gopher2021 · cited 5×
    “Training dataset size is also an important factor, but we do not plot capabilities against it because many language model families use a fixed number of training examples for all model sizes (Brown et al. 2020; Rae et al…”
    From this paper · §Emergent Abilities Definition
  • PaLM2022 · cited 5×
    “Training dataset size is also an important factor, but we do not plot capabilities against it because many language model families use a fixed number of training examples for all model sizes (Brown et al. 2020; Rae et al…”
    From this paper · §Emergent Abilities Definition
Show 8 more
  • Kaplan scaling laws2020 · cited 4×
    “In many cases, the effect of scale on performance can often be methodologically predicted via scaling laws—for example, scaling curves for cross-entropy loss have been shown to empirically span more than seven orders of…”
    From this paper · §Introduction
  • FLAN2021 · cited 4×
    “Consider the nascent direction of enabling language models to follow natural language instructions describing a task (Wei et al. 2022a; Sanh et al. 2022; Ouyang et al. 2022, inter alia).”
    From this paper · §Discussion
  • T02021 · cited 4×
    “Consider the nascent direction of enabling language models to follow natural language instructions describing a task (Wei et al. 2022a; Sanh et al. 2022; Ouyang et al. 2022, inter alia).”
    From this paper · §Discussion
  • BIG-bench2022 · cited 4×
    “Figure 2A–D depicts four emergent few-shot prompted tasks from BIG-Bench, a crowd-sourced suite of over 200 benchmarks for language model evaluation (BIG-Bench 2022).”
    From this paper · §Few-Shot Prompted Tasks
  • InstructGPT2022 · cited 3×
    “Consider the nascent direction of enabling language models to follow natural language instructions describing a task (Wei et al. 2022a; Sanh et al. 2022; Ouyang et al. 2022, inter alia).”
    From this paper · §Discussion
  • Flamingo2022 · cited 3×
    “As a few examples, GPT-3 175B achieved new state of the art on the TriviaQA and PiQA question-answering benchmarks (Brown et al. 2020); PaLM 540B achieved new state of the art on three arithmetic reasoning benchmarks (Ch…”
    From this paper · §Discussion
  • Switch Transformer2021 · cited 2×
    “For example, Chinchilla (Hoffmann et al. 2022) has one-fourth as many parameters as Gopher (Rae et al. 2021) but uses similar training compute; and sparse mixture-of-expert models have more parameters per training/infere…”
    From this paper · §Emergent Abilities Definition
  • GLaM2021 · cited 2×
    “For example, Chinchilla (Hoffmann et al. 2022) has one-fourth as many parameters as Gopher (Rae et al. 2021) but uses similar training compute; and sparse mixture-of-expert models have more parameters per training/infere…”
    From this paper · §Emergent Abilities Definition

Led to

  • U-PaLM2022 · cited 7×
    “To this end, large language models not only continue to improve as we scale in terms of data or computational budget (Hoffmann et al. 2022; Kaplan et al. 2020) but also acquire new abilities (Wei et al. 2022a).”
    From U-PaLM · §Related Work
  • EVA2022 · cited 2×
    “With further scaling on compute, data, and model sizes, PLMs have led to not only continuous performance improvements t5; kaplan2020scalingLM; rae2021gopher, but also a surprising emergence of in-context learning capabil…”
    From EVA · §Introduction
Abstract

Scaling up language models has been shown to predictably improve performance and sample efficiency on a wide range of downstream tasks. This paper instead discusses an unpredictable phenomenon that we refer to as emergent abilities of large language models. We consider an ability to be emergent if it is not present in smaller models but is present in larger models. Thus, emergent abilities cannot be predicted simply by extrapolating the performance of smaller models. The existence of such emergence implies that additional scaling could further expand the range of capabilities of language models.