PaLM: Scaling Language Modeling with Pathways
Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model PaLM.
Also cited · not yet reviewed (21)
- GLU variants2020 · cited 2×, 2 in Method“SwiGLU Activation – We use SwiGLU activations (Swish(xW)⋅xV\textrm{Swish}(xW)\cdot xV) for the MLP intermediate activations because they have been shown to significantly increase quality compared to standard ReLU, GeL…”From this paper · §Model Architecture
- GPT-32020 · cited 11×, 1 in Method“Because the weight initialization is proportional to 1/n1/\sqrt{n}, the effect of this is similar to the manual scaling down of Adam learning rate as in Brown et al. 2020.”From this paper · §Training Setup
- Transformer2017 · cited 3×, 1 in Method“PaLM uses a standard Transformer model architecture (Vaswani et al. 2017) in a decoder-only setup (i.e., each timestep can only attend to itself and past timesteps), with the following modifications:”From this paper · §Model Architecture
- SentencePiece2018 · cited 1×, 1 in Method“Vocabulary – We use a SentencePiece (Kudo & Richardson 2018a) vocabulary with 256k tokens, which was chosen to support the large number of languages in the training corpus without excess tokenization.”From this paper · §Model Architecture
Show 17 more
- RoFormer (RoPE)2021 · cited 1×, 1 in Method“RoPE Embeddings – We use RoPE embeddings (Su et al. 2021) rather than absolute or relative position embeddings, since RoPE embeddings have been shown to have better performance on long sequence lengths.”From this paper · §Model Architecture
- GLaM2021 · cited 12דThe most powerful of these post-GPT-3 models are GLaM (Du et al. 2021), Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), Megatron–Turing NLG (Smith et al. 2022), and LaMDA (Thoppilan et al. 2022), all of whic…”From this paper · §Introduction
- Gopher2021 · cited 11דThe most powerful of these post-GPT-3 models are GLaM (Du et al. 2021), Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), Megatron–Turing NLG (Smith et al. 2022), and LaMDA (Thoppilan et al. 2022), all of whic…”From this paper · §Introduction
- Codex2021 · cited 10דSecond, we compare to the early Codex model 12B described in Chen et al. 2021, which reports results only on the HumanEval dataset.”From this paper · §Evaluation
- LaMDA2022 · cited 8דThis dataset is based on the datasets used to train LaMDA (Thoppilan et al. 2022) and GLaM (Du et al. 2021).”From this paper · §Training Dataset
- Megatron-Turing NLG2022 · cited 8דThe most powerful of these post-GPT-3 models are GLaM (Du et al. 2021), Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), Megatron–Turing NLG (Smith et al. 2022), and LaMDA (Thoppilan et al. 2022), all of whic…”From this paper · §Introduction
- T52019 · cited 7דOn SuperGLUE, we compare with state-of-the-art models such as T5-11B (Raffel et al. 2020) and ST-MoE-32B (Zoph et al. 2022) and show that PaLM obtains competitive close-to-SOTA performance.”From this paper · §Evaluation
- GSM8K verifiers2021 · cited 7דSeveral recent papers have shown that large language models can achieve significant accuracy improvements by generating intermediate reasoning steps before generating the final answer (Nye et al. 2021; Cobbe et al. 2021;…”From this paper · §Evaluation
- FLAN2021 · cited 5דAny model that uses finetuning or multi-task adaptation (Wei et al. 2022a, Sanh et al. 2021) is not included in the table.”From this paper · §Evaluation
- Chinchilla2022 · cited 5דThe most powerful of these post-GPT-3 models are GLaM (Du et al. 2021), Gopher (Rae et al. 2021), Chinchilla (Hoffmann et al. 2022), Megatron–Turing NLG (Smith et al. 2022), and LaMDA (Thoppilan et al. 2022), all of whic…”From this paper · §Introduction
- Kaplan scaling laws2020 · cited 4דNote that scaling PaLM from 62B to 540B results in several drastic jumps in BLEU scores that do not follow the “power law” rule of thumb (Kaplan et al. 2020) projected from scaling the model size 8B to 62B.”From this paper · §Evaluation
- BERT2018 · cited 3דBroadly, language modeling refers to approaches for predicting either the next token in a sequence or for predicting masked spans (Devlin et al. 2019; Raffel et al. 2020).”From this paper · §Related Work
- GPipe2018 · cited 3דTherefore, techniques have arisen for splitting model tensors across accelerators (Shazeer et al. 2018) or alternatively separating layers of the models across accelerators and then pipe-lining activations between the st…”From this paper · §Related Work
- GShard2020 · cited 2דMany other works aim to increase of the scale of models, while limiting communication overheads (Rajbhandari et al. 2020; Lepikhin et al. 2020; Li et al. 2020; Rasley et al. 2020; Rajbhandari et al. 2021; Ren et al. 2021…”From this paper · §Related Work
- Switch Transformer2021 · cited 2דScale has continued to increase after GPT-3, evidenced by the succession of the 178B parameter Jurassic-1 (Lieber et al. 2021), the 280B parameter Gopher model (Rae et al. 2021), the 530B Megatron-Turing NLG (Smith et al…”From this paper · §Related Work
- T02021 · cited 2דAny model that uses finetuning or multi-task adaptation (Wei et al. 2022a, Sanh et al. 2021) is not included in the table.”From this paper · §Evaluation
- BIG-bench2022 · cited 2דIn section 6.2 we additionally highlight breakthrough performance on BIG-bench (BIG-bench collaboration 2021), a recently released suite of 150150+ new language understanding and generation tasks, many of which are extre…”From this paper · §Introduction
Led to
- GPT-NeoX-20B2022 · cited 2דA consequence of this has been an abundance of research focusing on scaling Transformer models up to ever-larger scales, resulting in dense models that surpass 500B parameters (Smith et al. 2022; Chowdhery et al. 2022),…”From GPT-NeoX-20B · §Introduction
- OPT2022 · cited 10דFollowing PaLM Chowdhery et al. 2022, we sample 25 generations of 20 tokens using nucleus sampling Holtzman et al. 2020 (p=0.9p=0.9) for each of 10,00010,000 randomly sampled prompts from RTP, and report mean toxicity pr…”From OPT · §Bias & Toxicity Evaluations
- BIG-bench2022 · cited 4×, 2 in Method“Recent results on large language models suggest that this brittleness to question phrasing may improve with further increases in scale (Chowdhery et al. 2022).”From BIG-bench · §Behavior of language models and human raters on BIG-bench
- Emergent abilities2022 · cited 5דTraining dataset size is also an important factor, but we do not plot capabilities against it because many language model families use a fixed number of training examples for all model sizes (Brown et al. 2020; Rae et al…”From Emergent abilities · §Emergent Abilities Definition
- U-PaLM2022 · cited 6דScaling and improving large language models is one of the most impactful research areas in modern artificial intelligence (Chowdhery et al. 2022).”From U-PaLM · §Related Work
- Flan-T5 / Flan-PaLM2022 · cited 5דAs an extended version of the PaLM 62B model, we instruction-finetune cont-PaLM, which is a 62B PaLM-model initialized from PaLM-62B and then pretrained for 500B more tokens (Chowdhery et al. 2022).”From Flan-T5 / Flan-PaLM · §unknown section
Abstract
Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model PaLM. We trained PaLM on 6144 TPU v4 chips using Pathways, a new ML system which enables highly efficient training across multiple TPU Pods. We demonstrate continued benefits of scaling by achieving state-of-the-art few-shot learning results on hundreds of language understanding and generation benchmarks. On a number of these tasks, PaLM 540B achieves breakthrough performance, outperforming the finetuned state-of-the-art on a suite of multi-step reasoning tasks, and outperforming average human performance on the recently released BIG-bench benchmark. A significant number of BIG-bench tasks showed discontinuous improvements from model scale, meaning that performance steeply increased as we scaled to our largest model. PaLM also has strong capabilities in multilingual tasks and source code generation, which we demonstrate on a wide array of benchmarks. We additionally provide a comprehensive analysis on bias and toxicity, and study the extent of training data memorization with respect to model scale. Finally, we discuss the ethical considerations related to large language models and discuss potential mitigation strategies.