Paper Lineage
Esc
MethodApr 2022arXiv 2204.02311cs.CL

PaLM: Scaling Language Modeling with Pathways

Aakanksha Chowdhery, Sharan Narang, Jacob Devlin and 64 others

Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model PaLM.

From the abstract

Built on

21 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (21)

  • GLU variants2020 · cited 2×, 2 in Method
    “SwiGLU Activation – We use SwiGLU activations (Swish​(x​W)⋅x​V\textrm{Swish}(xW)\cdot xV) for the MLP intermediate activations because they have been shown to significantly increase quality compared to standard ReLU, GeL…”
    From this paper · §Model Architecture
  • GPT-32020 · cited 11×, 1 in Method
    “Because the weight initialization is proportional to 1/n1/\sqrt{n}, the effect of this is similar to the manual scaling down of Adam learning rate as in Brown et al. 2020.”
    From this paper · §Training Setup
  • Transformer2017 · cited 3×, 1 in Method
    “PaLM uses a standard Transformer model architecture (Vaswani et al. 2017) in a decoder-only setup (i.e., each timestep can only attend to itself and past timesteps), with the following modifications:”
    From this paper · §Model Architecture
  • SentencePiece2018 · cited 1×, 1 in Method
    “Vocabulary – We use a SentencePiece (Kudo & Richardson 2018a) vocabulary with 256k tokens, which was chosen to support the large number of languages in the training corpus without excess tokenization.”
    From this paper · §Model Architecture
Show 17 more
  • RoFormer (RoPE)2021 · cited 1×, 1 in Method
    “RoPE Embeddings – We use RoPE embeddings (Su et al. 2021) rather than absolute or relative position embeddings, since RoPE embeddings have been shown to have better performance on long sequence lengths.”
    From this paper · §Model Architecture
  • GLaM2021 · cited 12×
    “The most powerful of these post-GPT-3 models are GLaM (Du et al. 2021), Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), Megatron–Turing NLG (Smith et al. 2022), and LaMDA (Thoppilan et al. 2022), all of whic…”
    From this paper · §Introduction
  • Gopher2021 · cited 11×
    “The most powerful of these post-GPT-3 models are GLaM (Du et al. 2021), Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), Megatron–Turing NLG (Smith et al. 2022), and LaMDA (Thoppilan et al. 2022), all of whic…”
    From this paper · §Introduction
  • Codex2021 · cited 10×
    “Second, we compare to the early Codex model 12B described in Chen et al. 2021, which reports results only on the HumanEval dataset.”
    From this paper · §Evaluation
  • LaMDA2022 · cited 8×
    “This dataset is based on the datasets used to train LaMDA (Thoppilan et al. 2022) and GLaM (Du et al. 2021).”
    From this paper · §Training Dataset
  • Megatron-Turing NLG2022 · cited 8×
    “The most powerful of these post-GPT-3 models are GLaM (Du et al. 2021), Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), Megatron–Turing NLG (Smith et al. 2022), and LaMDA (Thoppilan et al. 2022), all of whic…”
    From this paper · §Introduction
  • T52019 · cited 7×
    “On SuperGLUE, we compare with state-of-the-art models such as T5-11B (Raffel et al. 2020) and ST-MoE-32B (Zoph et al. 2022) and show that PaLM obtains competitive close-to-SOTA performance.”
    From this paper · §Evaluation
  • GSM8K verifiers2021 · cited 7×
    “Several recent papers have shown that large language models can achieve significant accuracy improvements by generating intermediate reasoning steps before generating the final answer (Nye et al. 2021; Cobbe et al. 2021;…”
    From this paper · §Evaluation
  • FLAN2021 · cited 5×
    “Any model that uses finetuning or multi-task adaptation (Wei et al. 2022a, Sanh et al. 2021) is not included in the table.”
    From this paper · §Evaluation
  • Chinchilla2022 · cited 5×
    “The most powerful of these post-GPT-3 models are GLaM (Du et al. 2021), Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), Megatron–Turing NLG (Smith et al. 2022), and LaMDA (Thoppilan et al. 2022), all of whic…”
    From this paper · §Introduction
  • Kaplan scaling laws2020 · cited 4×
    “Note that scaling PaLM from 62B to 540B results in several drastic jumps in BLEU scores that do not follow the “power law” rule of thumb (Kaplan et al. 2020) projected from scaling the model size 8B to 62B.”
    From this paper · §Evaluation
  • BERT2018 · cited 3×
    “Broadly, language modeling refers to approaches for predicting either the next token in a sequence or for predicting masked spans (Devlin et al. 2019; Raffel et al. 2020).”
    From this paper · §Related Work
  • GPipe2018 · cited 3×
    “Therefore, techniques have arisen for splitting model tensors across accelerators (Shazeer et al. 2018) or alternatively separating layers of the models across accelerators and then pipe-lining activations between the st…”
    From this paper · §Related Work
  • GShard2020 · cited 2×
    “Many other works aim to increase of the scale of models, while limiting communication overheads (Rajbhandari et al. 2020; Lepikhin et al. 2020; Li et al. 2020; Rasley et al. 2020; Rajbhandari et al. 2021; Ren et al. 2021…”
    From this paper · §Related Work
  • Switch Transformer2021 · cited 2×
    “Scale has continued to increase after GPT-3, evidenced by the succession of the 178B parameter Jurassic-1 (Lieber et al. 2021), the 280B parameter Gopher model (Rae et al. 2021), the 530B Megatron-Turing NLG (Smith et al…”
    From this paper · §Related Work
  • T02021 · cited 2×
    “Any model that uses finetuning or multi-task adaptation (Wei et al. 2022a, Sanh et al. 2021) is not included in the table.”
    From this paper · §Evaluation
  • BIG-bench2022 · cited 2×
    “In section 6.2 we additionally highlight breakthrough performance on BIG-bench (BIG-bench collaboration 2021), a recently released suite of 150150+ new language understanding and generation tasks, many of which are extre…”
    From this paper · §Introduction

Led to

  • GPT-NeoX-20B2022 · cited 2×
    “A consequence of this has been an abundance of research focusing on scaling Transformer models up to ever-larger scales, resulting in dense models that surpass 500B parameters (Smith et al. 2022; Chowdhery et al. 2022),…”
    From GPT-NeoX-20B · §Introduction
  • OPT2022 · cited 10×
    “Following PaLM Chowdhery et al. 2022, we sample 25 generations of 20 tokens using nucleus sampling Holtzman et al. 2020 (p=0.9p=0.9) for each of 10,00010,000 randomly sampled prompts from RTP, and report mean toxicity pr…”
    From OPT · §Bias & Toxicity Evaluations
  • BIG-bench2022 · cited 4×, 2 in Method
    “Recent results on large language models suggest that this brittleness to question phrasing may improve with further increases in scale (Chowdhery et al. 2022).”
    From BIG-bench · §Behavior of language models and human raters on BIG-bench
  • Emergent abilities2022 · cited 5×
    “Training dataset size is also an important factor, but we do not plot capabilities against it because many language model families use a fixed number of training examples for all model sizes (Brown et al. 2020; Rae et al…”
    From Emergent abilities · §Emergent Abilities Definition
  • U-PaLM2022 · cited 6×
    “Scaling and improving large language models is one of the most impactful research areas in modern artificial intelligence (Chowdhery et al. 2022).”
    From U-PaLM · §Related Work
  • Flan-T5 / Flan-PaLM2022 · cited 5×
    “As an extended version of the PaLM 62B model, we instruction-finetune cont-PaLM, which is a 62B PaLM-model initialized from PaLM-62B and then pretrained for 500B more tokens (Chowdhery et al. 2022).”
    From Flan-T5 / Flan-PaLM · §unknown section
Abstract

Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model PaLM. We trained PaLM on 6144 TPU v4 chips using Pathways, a new ML system which enables highly efficient training across multiple TPU Pods. We demonstrate continued benefits of scaling by achieving state-of-the-art few-shot learning results on hundreds of language understanding and generation benchmarks. On a number of these tasks, PaLM 540B achieves breakthrough performance, outperforming the finetuned state-of-the-art on a suite of multi-step reasoning tasks, and outperforming average human performance on the recently released BIG-bench benchmark. A significant number of BIG-bench tasks showed discontinuous improvements from model scale, meaning that performance steeply increased as we scaled to our largest model. PaLM also has strong capabilities in multilingual tasks and source code generation, which we demonstrate on a wide array of benchmarks. We additionally provide a comprehensive analysis on bias and toxicity, and study the extent of training data memorization with respect to model scale. Finally, we discuss the ethical considerations related to large language models and discuss potential mitigation strategies.