Emergent Abilities of Large Language Models
Scaling up language models has been shown to predictably improve performance and sample efficiency on a wide range of downstream tasks. This paper instead discusses an unpredictable phenomenon that we refer to as emergent abilities of large language models.
Also cited · not yet reviewed (12)
- GPT-32020 · cited 9דIt is now well-known that increasing the scale of language models (e.g., training compute, model parameters, etc.) can lead to better performance and sample efficiency on a range of downstream NLP tasks (Devlin et al. 20…”From this paper · §Introduction
- Chinchilla2022 · cited 6דIn many cases, the effect of scale on performance can often be methodologically predicted via scaling laws—for example, scaling curves for cross-entropy loss have been shown to empirically span more than seven orders of…”From this paper · §Introduction
- Gopher2021 · cited 5דTraining dataset size is also an important factor, but we do not plot capabilities against it because many language model families use a fixed number of training examples for all model sizes (Brown et al. 2020; Rae et al…”From this paper · §Emergent Abilities Definition
- PaLM2022 · cited 5דTraining dataset size is also an important factor, but we do not plot capabilities against it because many language model families use a fixed number of training examples for all model sizes (Brown et al. 2020; Rae et al…”From this paper · §Emergent Abilities Definition
Show 8 more
- Kaplan scaling laws2020 · cited 4דIn many cases, the effect of scale on performance can often be methodologically predicted via scaling laws—for example, scaling curves for cross-entropy loss have been shown to empirically span more than seven orders of…”From this paper · §Introduction
- FLAN2021 · cited 4דConsider the nascent direction of enabling language models to follow natural language instructions describing a task (Wei et al. 2022a; Sanh et al. 2022; Ouyang et al. 2022, inter alia).”From this paper · §Discussion
- T02021 · cited 4דConsider the nascent direction of enabling language models to follow natural language instructions describing a task (Wei et al. 2022a; Sanh et al. 2022; Ouyang et al. 2022, inter alia).”From this paper · §Discussion
- BIG-bench2022 · cited 4דFigure 2A–D depicts four emergent few-shot prompted tasks from BIG-Bench, a crowd-sourced suite of over 200 benchmarks for language model evaluation (BIG-Bench 2022).”From this paper · §Few-Shot Prompted Tasks
- InstructGPT2022 · cited 3דConsider the nascent direction of enabling language models to follow natural language instructions describing a task (Wei et al. 2022a; Sanh et al. 2022; Ouyang et al. 2022, inter alia).”From this paper · §Discussion
- Flamingo2022 · cited 3דAs a few examples, GPT-3 175B achieved new state of the art on the TriviaQA and PiQA question-answering benchmarks (Brown et al. 2020); PaLM 540B achieved new state of the art on three arithmetic reasoning benchmarks (Ch…”From this paper · §Discussion
- Switch Transformer2021 · cited 2דFor example, Chinchilla (Hoffmann et al. 2022) has one-fourth as many parameters as Gopher (Rae et al. 2021) but uses similar training compute; and sparse mixture-of-expert models have more parameters per training/infere…”From this paper · §Emergent Abilities Definition
- GLaM2021 · cited 2דFor example, Chinchilla (Hoffmann et al. 2022) has one-fourth as many parameters as Gopher (Rae et al. 2021) but uses similar training compute; and sparse mixture-of-expert models have more parameters per training/infere…”From this paper · §Emergent Abilities Definition
Led to
- U-PaLM2022 · cited 7דTo this end, large language models not only continue to improve as we scale in terms of data or computational budget (Hoffmann et al. 2022; Kaplan et al. 2020) but also acquire new abilities (Wei et al. 2022a).”From U-PaLM · §Related Work
- EVA2022 · cited 2דWith further scaling on compute, data, and model sizes, PLMs have led to not only continuous performance improvements t5; kaplan2020scalingLM; rae2021gopher, but also a surprising emergence of in-context learning capabil…”From EVA · §Introduction
Abstract
Scaling up language models has been shown to predictably improve performance and sample efficiency on a wide range of downstream tasks. This paper instead discusses an unpredictable phenomenon that we refer to as emergent abilities of large language models. We consider an ability to be emergent if it is not present in smaller models but is present in larger models. Thus, emergent abilities cannot be predicted simply by extrapolating the performance of smaller models. The existence of such emergence implies that additional scaling could further expand the range of capabilities of language models.