Paper Lineage
Esc
MethodMay 2020arXiv 2005.14165cs.CL

Language Models are Few-Shot Learners

Tom B. Brown, Benjamin Mann, Nick Ryder and 28 others

Recent work has demonstrated substantial gains on many NLP tasks and benchmarks by pre-training on a large corpus of text followed by fine-tuning on a specific task. While typically task-agnostic in architecture, this method still requires task-specific fine-tuning datasets of thousands or tens of thousands of examples.

From the abstract

Built on

12 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (12)

  • Kaplan scaling laws2020 · cited 7×, 4 in Method
    “Previous work [57] suggests that with enough training data, scaling of validation loss should be approximately a smooth power law as a function of size; training models of many different sizes allows us to test this hypo…”
    From this paper · §Approach
  • T52019 · cited 10×, 3 in Method
    “On tasks that involve binary classification, we give the options more semantically meaningful names (e.g. “True” or “False” rather than 0 or 1) and then treat the task like multiple choice; we also sometimes frame the ta…”
    From this paper · §Approach
  • BERT2018 · cited 4×
    “Work in this vein has successively increased model size: 213 million parameters [134] in the original paper, 300 million parameters [20], 1.5 billion parameters [117], 8 billion parameters [125], 11 billion parameters [1…”
    From this paper · §Related Work
  • “Recent efforts include [116, 115], which fine-tuned an 11 billion parameter language model, and [33], which focused on attending over a large corpus of data at test time.”
    From this paper · §Related Work
Show 8 more
  • Megatron-LM2019 · cited 3×
    “Work in this vein has successively increased model size: 213 million parameters [134] in the original paper, 300 million parameters [20], 1.5 billion parameters [117], 8 billion parameters [125], 11 billion parameters [1…”
    From this paper · §Related Work
  • ALBERT2019 · cited 3×
    “This last paradigm has led to substantial progress on many challenging NLP tasks such as reading comprehension, question answering, textual entailment, and many others, and has continued to advance based on new architect…”
    From this paper · §Introduction
  • Knowledge Distillation2015 · cited 2×
    “This approach includes ALBERT [62] as well as general [44] and task-specific [121, 52, 59] approaches to distillation of language models.”
    From this paper · §Related Work
  • Transformer2017 · cited 2×
    “Work in this vein has successively increased model size: 213 million parameters [134] in the original paper, 300 million parameters [20], 1.5 billion parameters [117], 8 billion parameters [125], 11 billion parameters [1…”
    From this paper · §Related Work
  • CoQA2018 · cited 2×
    “As fine-tuned language models have neared human performance on many standard benchmark tasks, considerable effort has been devoted to constructing more difficult or open-ended tasks, including question answering [58, 47,…”
    From this paper · §Related Work
  • XLNet2019 · cited 2×
    “This last paradigm has led to substantial progress on many challenging NLP tasks such as reading comprehension, question answering, textual entailment, and many others, and has continued to advance based on new architect…”
    From this paper · §Introduction
  • RoBERTa2019 · cited 2×
    “This last paradigm has led to substantial progress on many challenging NLP tasks such as reading comprehension, question answering, textual entailment, and many others, and has continued to advance based on new architect…”
    From this paper · §Introduction
  • RLHF for LMs (Ziegler)2019 · cited 2×
    “One direction for future work might be attempting to generate a broader set of explicit tasks for multi-task learning, for example through procedural generation [128], human interaction [144], or active learning [80].”
    From this paper · §Related Work

Led to

  • VirTex2020 · cited 2×, 1 in Method
    “This involves training massive language models – either unidirectional [77] or bidirectional [78, 79, 80, 81], for predicting tokens one by one.”
    From VirTex · §Method
  • GShard2020 · cited 3×
    “Scaling neural networks brings dramatic quality gains over a wide array of machine learning problems [1, 2, 3, 4, 5, 6].”
    From GShard · §Introduction
  • “We follow the works of [48, 4] and use large pretrained GPT-3 models with as many as 6.7 billion parameters.”
    From Learning to summarize from human feedbac · §unknown section
  • MMLU2020 · cited 4×
    “However, larger pretrained models like GPT-3 (Brown et al. 2020) have made it possible to achieve competitive performance without fine-tuning by using few-shot learning, which removes the need for a large fine-tuning set…”
    From MMLU · §Related Work
  • ViT2020 · cited 2×
    “Large Transformer-based models are often pre-trained on large corpora and then fine-tuned for the task at hand: BERT (Devlin et al. 2019) uses a denoising self-supervised pre-training task, while the GPT line of work use…”
    From ViT · §Related Work
  • The Pile2020 · cited 15×, 1 in Method
    “Following Brown et al. 2020, we increase the weights of higher quality components, with certain high-quality datasets such as Wikipedia being seen up to 3 times (“epochs”) for each full epoch over the Pile.”
    From The Pile · §The Pile Datasets
  • Prefix-Tuning2021 · cited 3×
    “GPT-3 Brown et al. 2020 uses manually designed prompts to adapt its generation for different tasks, and this framework is termed in-context learning.”
    From Prefix-Tuning · §Related Work
  • Switch Transformer2021 · cited 4×, 1 in Method
    “It is common to mix both model and data parallelism for large scale models, which was done in the largest T5 models (Raffel et al. 2019; Xue et al. 2020) and in GPT-3 (Brown et al. 2020).”
    From Switch Transformer · §Designing Models with Data, Model, and Expert-Parallelism
  • CLIP2021 · cited 4×
    “Similar to the “prompt engineering” discussion around GPT-3 (Brown et al. 2020; Gao et al. 2020), we have also observed that zero-shot performance can be significantly improved by customizing the prompt text to each task…”
    From CLIP · §Experiments
  • Prompt tuning2021 · cited 2×
    “More recently, Brown et al. 2020 showed that prompt design (or “priming”) is surprisingly effective at modulating a frozen GPT-3 model’s behavior through text prompts.”
    From Prompt tuning · §Introduction
  • Natural Instructions2021 · cited 3×, 2 in Method
    “We build models using pre-trained LMs with encoder-decoder architectures BART Lewis et al. 2019 for fine-tuning and GPT3 Brown et al. 2020 for few-shot experiments.”
    From Natural Instructions · §Problem Setup and Models
  • True few-shot learning2021 · cited 11×, 2 in Method
    “Recent work does not assume access to data from other distributions, performing few-shot learning using only a few examples from a single distribution to update a pretrained LM [2, 12].”
    From True few-shot learning · §Can We Do Model Selection in Few-Shot Learning?
  • Scaling ViTs (ViT-G)2021 · cited 3×
    “The few-shot transfer evaluation protocol has also been adopted by previous large-scale pre-training efforts in NLP domain gpt3.”
    From Scaling ViTs (ViT-G) · §Introduction
  • Frozen2021 · cited 4×
    “Auto-regressive transformers have been shown to be very impressive models of natural language [40].”
    From Frozen · §Introduction
  • Codex2021 · cited 4×
    “Following the success of large natural language models (Devlin et al. 2018; Radford et al. 2019; Liu et al. 2019; Raffel et al. 2020; Brown et al. 2020) large scale Transformers have also been applied towards program syn…”
    From Codex · §Related Work
  • SimVLM2021 · cited 4×
    “Self-supervised textual representation learning (Devlin et al. 2018; Radford et al. 2018; Radford et al. 2019; Liu et al. 2019; Yang et al. 2019; Raffel et al. 2019; Brown et al. 2020) based on Transformers (Vaswani et a…”
    From SimVLM · §Introduction
  • FLAN2021 · cited 8×
    “Language models (LMs) at scale, such as GPT-3 (Brown et al. 2020), have been shown to perform few-shot learning remarkably well.”
    From FLAN · §Introduction
  • “We use pretrained transformer language models (Vaswani et al., 2017) from the GPT-3 family (Brown et al., 2020), which take 2048 tokens of context.”
    From Recursive book summarization · §unknown section
  • T02021 · cited 13×, 1 in Method
    “Finally, in explaining the success of prompts, the leading hypothesis is that models learn to understand the prompts as task instructions which help them generalize to held-out tasks (Wei et al. 2021; Mishra et al. 2021;…”
    From T0 · §Related Work
  • GSM8K verifiers2021 · cited 2×, 1 in Method
    “Finetuning, our baseline method, uses the same language modeling objective as the generative pretraining in GPT-3 (Brown et al. 2020).”
    From GSM8K verifiers · §Methods
  • MAE2021 · cited 6×
    “The solutions, based on autoregressive language modeling in GPT Radford2018; Radford2019; Brown2020 and masked autoencoding in BERT Devlin2019, are conceptually simple: they remove a portion of the data and learn to pred…”
    From MAE · §Introduction
  • Swin V22021 · cited 4×
    “It significantly improves a model’s performance on language tasks devlin2018bert; radford2019language; raffel2019t5; Turing-17B; fedus2021switch; Megatron-Turing-530B and the model demonstrates amazing few-shot capabilit…”
    From Swin V2 · §Introduction
  • Gopher2021 · cited 7×
    “The baseline model includes LLMs such as GPT-3 (175B parameters) (Brown et al. 2020), Jurassic-1 (Lieber et al. 2021) (178B parameters), and Megatron-Turing NLG (530B parameters) (Kharya and Alvi 2021); the exact baselin…”
    From Gopher · §Results
  • ViTCAP2021 · cited 2×
    “The transformer architecture and its instantiations (e.g., BERT devlin2018bert, GPT brown2020language) are well-known for their remarkable performances on natural language processing tasks, which are mostly attributed to…”
    From ViTCAP · §ViTCAP
  • GLaM2021 · cited 13×, 3 in Method
    “GPT-3 (Brown et al. 2020) and related work (Shoeybi et al. 2019; Lieber et al. 2021; Wei et al. 2021) demonstrated that scaling up language models greatly improves task-agnostic, few-shot performance.”
    From GLaM · §Related Work
  • Fairseq MoE LMs2021 · cited 25×, 19 in Method
    “We train autoregressive (decoder-only) transformer models that roughly match the sizes and architecture explored in Brown et al. 2020.”
    From Fairseq MoE LMs · §Experimental Setup
  • LaMDA2022 · cited 7×
    “Similar to the concept of prompts in GPT-3 [12], we precondition LaMDA on a few turns of application-specific dialog to adapt LaMDA to the target applications.”
    From LaMDA · §Introduction
  • Megatron-Turing NLG2022 · cited 23×, 7 in Method
    “As prior work has found (e.g., [9]), the quality of unfiltered Common Crawl data is lower than that of curated datasets and steps should be taken to increase the average quality of data selected from Common Crawl for LM…”
    From Megatron-Turing NLG · §Training Dataset and Model Configuration
  • InstructGPT2022 · cited 4×, 2 in Method
    “We start with the GPT-3 pretrained language models from Brown et al., 2020.”
    From InstructGPT · §Methods and experimental details
  • Chinchilla2022 · cited 3×
    “Following Kaplan et al. 2020 and the training setup of GPT-3 (Brown et al. 2020), many of the recently trained large models have been trained for approximately 300 billion tokens (Table 1), in line with the approach of p…”
    From Chinchilla · §Introduction
  • PaLM2022 · cited 11×, 1 in Method
    “Because the weight initialization is proportional to 1/n1/\sqrt{n}, the effect of this is similar to the manual scaling down of Adam learning rate as in Brown et al. 2020.”
    From PaLM · §Training Setup
  • GPT-NeoX-20B2022 · cited 10×, 2 in Method
    “GPT-NeoX-20B is an autoregressive transformer decoder model whose architecture largely follows that of GPT-3 (Brown et al. 2020), with a few notable deviations described below.”
    From GPT-NeoX-20B · §Model Design and Implementation
  • Super-NaturalInstructions2022 · cited 2×
    “Additionally, we evaluate GPT-3 Brown et al. 2020, a 175175B-parameter autoregressive LM that has shown remarkable ability in following demonstrations provided in its prompt.”
    From Super-NaturalInstructions · §Benchmarking Cross-Task Generalization with Sup-NatInst
  • Flamingo2022 · cited 6×, 1 in Method
    “We evaluate the ability of our models to rapidly adapt to new tasks using in-context learning, analogously to GPT-3 [11], by interleaving support example pairs in the form of (i​m​a​g​e,t​e​x​t)(image,text) or (v​i​d​e​o…”
    From Flamingo · §Approach
  • OPT2022 · cited 15×, 2 in Method
    “In the interest of transparency, and to reduce risk of training instabilities, our models and hyperparameters largely follow Brown et al. 2020, with variations in batch size mostly to obtain increased computational effic…”
    From OPT · §Method
  • BIG-bench2022 · cited 26×, 1 in Method
    “We use OpenAI GPT models corresponding to the GPT-3 model series in Brown et al. 2020.”
    From BIG-bench · §What is in BIG-bench?
  • Emergent abilities2022 · cited 9×
    “It is now well-known that increasing the scale of language models (e.g., training compute, model parameters, etc.) can lead to better performance and sample efficiency on a range of downstream NLP tasks (Devlin et al. 20…”
    From Emergent abilities · §Introduction
  • U-PaLM2022 · cited 4×
    “For evaluation, we use the average score of NLU and NLG tasks from the GPT-3 suite (Brown et al. 2020).”
    From U-PaLM · §Experiments
  • Flan-T5 / Flan-PaLM2022 · cited 6×
    “In natural language processing (NLP), pretrained language models have made significant progress towards this goal, as they can perform tasks given natural language descriptions (Brown et al. 2020, inter alia).”
    From Flan-T5 / Flan-PaLM · §unknown section
  • EVA2022 · cited 2×
    “Scaling up pre-trained language models (PLMs) liu2019roberta; gpt3; t5 has revolutionized natural language processing (NLP) in the past few years.”
    From EVA · §Introduction
Abstract

Recent work has demonstrated substantial gains on many NLP tasks and benchmarks by pre-training on a large corpus of text followed by fine-tuning on a specific task. While typically task-agnostic in architecture, this method still requires task-specific fine-tuning datasets of thousands or tens of thousands of examples. By contrast, humans can generally perform a new language task from only a few examples or from simple instructions - something which current NLP systems still largely struggle to do. Here we show that scaling up language models greatly improves task-agnostic, few-shot performance, sometimes even reaching competitiveness with prior state-of-the-art fine-tuning approaches. Specifically, we train GPT-3, an autoregressive language model with 175 billion parameters, 10x more than any previous non-sparse language model, and test its performance in the few-shot setting. For all tasks, GPT-3 is applied without any gradient updates or fine-tuning, with tasks and few-shot demonstrations specified purely via text interaction with the model. GPT-3 achieves strong performance on many NLP datasets, including translation, question-answering, and cloze tasks, as well as several tasks that require on-the-fly reasoning or domain adaptation, such as unscrambling words, using a novel word in a sentence, or performing 3-digit arithmetic. At the same time, we also identify some datasets where GPT-3's few-shot learning still struggles, as well as some datasets where GPT-3 faces methodological issues related to training on large web corpora. Finally, we find that GPT-3 can generate samples of news articles which human evaluators have difficulty distinguishing from articles written by humans. We discuss broader societal impacts of this finding and of GPT-3 in general.