Paper Lineage
Esc
AnalysisJan 2020arXiv 2001.08361cs.LG

Scaling Laws for Neural Language Models

Jared Kaplan, Sam McCandlish, Tom Henighan and 7 others

We study empirical scaling laws for language model performance on the cross-entropy loss. The loss scales as a power-law with model size, dataset size, and the amount of compute used for training, with some trends spanning more than seven orders of magnitude.

From the abstract

Built on

3 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (3)

  • Transformer2017 · cited 2×, 1 in Method
    “In this work we will empirically investigate the dependence of language modeling loss on all of these factors, focusing on the Transformer architecture [VSP+17, LSP+18].”
    From this paper · §Introduction
  • BookCorpus (books & movies)2015 · cited 1×, 1 in Method
    “We reserve 6.6×1086.6\times 10^{8} of these tokens for use as a test set, and we also test on similarly-prepared samples of Books Corpus [ZKZ+15], Common Crawl [Fou], English Wikipedia, and a collection of publicly-avail…”
    From this paper · §Background and Methods
  • EfficientNet2019 · cited 2×
    “EfficientNets [TL19] also appear to obey an approximate power-law relation between accuracy and model size.”
    From this paper · §Related Work

Led to

  • GPT-32020 · cited 7×, 4 in Method
    “Previous work [57] suggests that with enough training data, scaling of validation loss should be approximately a smooth power law as a function of size; training models of many different sizes allows us to test this hypo…”
    From GPT-3 · §Approach
  • GShard2020 · cited 4×
    “Scaling neural networks brings dramatic quality gains over a wide array of machine learning problems [1, 2, 3, 4, 5, 6].”
    From GShard · §Introduction
  • The Pile2020 · cited 2×, 1 in Method
    “Recent work has shown an increased focus on the empirical scaling laws of language models (Kaplan et al. 2020; Henighan et al. 2020).”
    From The Pile · §Benchmarking Language Models with the Pile
  • Switch Transformer2021 · cited 6×
    “Large scale training has been an effective path towards flexible and powerful neural language models (Radford et al. 2018; Kaplan et al. 2020; Brown et al. 2020).”
    From Switch Transformer · §Introduction
  • CLIP2021 · cited 2×
    “We study the scalability of CLIP by training a series of eight models spanning almost 2 orders of magnitude of compute and observe that transfer performance is a smoothly predictable function of compute (Hestness et al.…”
    From CLIP · §Introduction and Motivating Work
  • Scaling ViTs (ViT-G)2021 · cited 3×
    “Optimal scaling of Transformers in NLP was carefully studied in kaplan2020scaling, with the main conclusion that large models not only perform better, but do use large computational budgets more efficiently.”
    From Scaling ViTs (ViT-G) · §Introduction
  • LaMDA2022 · cited 3×
    “Our study of scaling laws with respect to model sizes is inspired by recent work on the scaling laws of neural language models [12, 13].”
    From LaMDA · §Related work
  • Chinchilla2022 · cited 12×
    “Following Kaplan et al. 2020 and the training setup of GPT-3 (Brown et al. 2020), many of the recently trained large models have been trained for approximately 300 billion tokens (Table 1), in line with the approach of p…”
    From Chinchilla · §Introduction
  • PaLM2022 · cited 4×
    “Note that scaling PaLM from 62B to 540B results in several drastic jumps in BLEU scores that do not follow the “power law” rule of thumb (Kaplan et al. 2020) projected from scaling the model size 8B to 62B.”
    From PaLM · §Evaluation
  • GPT-NeoX-20B2022 · cited 3×, 1 in Method
    “Our model has 20 billion parameters, of which 19.9 billion are “non-embedding” parameters that Kaplan et al. 2020 identify as the proper number to use for scaling laws analysis.”
    From GPT-NeoX-20B · §Model Design and Implementation
  • BIG-bench2022 · cited 2×
    “Their cross entropy on a test set scales as a power law with regards to model size, training data size, and the amount of compute used in training (Hestness et al. 2017; Hestness et al. 2019; Rosenfeld et al. 2019; Kapla…”
    From BIG-bench · §Introduction
  • Emergent abilities2022 · cited 4×
    “In many cases, the effect of scale on performance can often be methodologically predicted via scaling laws—for example, scaling curves for cross-entropy loss have been shown to empirically span more than seven orders of…”
    From Emergent abilities · §Introduction
  • U-PaLM2022 · cited 5×
    “To this end, large language models not only continue to improve as we scale in terms of data or computational budget (Hoffmann et al. 2022; Kaplan et al. 2020) but also acquire new abilities (Wei et al. 2022a).”
    From U-PaLM · §Related Work
Abstract

We study empirical scaling laws for language model performance on the cross-entropy loss. The loss scales as a power-law with model size, dataset size, and the amount of compute used for training, with some trends spanning more than seven orders of magnitude. Other architectural details such as network width or depth have minimal effects within a wide range. Simple equations govern the dependence of overfitting on model/dataset size and the dependence of training speed on model size. These relationships allow us to determine the optimal allocation of a fixed compute budget. Larger models are significantly more sample-efficient, such that optimally compute-efficient training involves training very large models on a relatively modest amount of data and stopping significantly before convergence.