OPT: Open Pre-trained Transformer Language Models
Large language models, which are often trained for hundreds of thousands of compute days, have shown remarkable capabilities for zero- and few-shot learning. Given their computational cost, these models are difficult to replicate without significant capital.
Also cited · not yet reviewed (14)
- GPT-32020 · cited 15×, 2 in Method“In the interest of transparency, and to reduce risk of training instabilities, our models and hyperparameters largely follow Brown et al. 2020, with variations in batch size mostly to obtain increased computational effic…”From this paper · §Method
- RoBERTa2019 · cited 3×, 2 in Method“The pre-training corpus contains a concatenation of datasets used in RoBERTa Liu et al. 2019b, the Pile Gao et al. 2021a, and PushShift.io Reddit Baumgartner et al. 2020; Roller et al. 2021.”From this paper · §Method
- The Pile2020 · cited 2×, 2 in Method“The pre-training corpus contains a concatenation of datasets used in RoBERTa Liu et al. 2019b, the Pile Gao et al. 2021a, and PushShift.io Reddit Baumgartner et al. 2020; Roller et al. 2021.”From this paper · §Method
- Fairseq MoE LMs2021 · cited 3×, 1 in Method“Following Lieber et al. 2021 and Artetxe et al. 2021, we use StereoSet Nadeem et al. 2021 to measure stereotypical bias across 4 categories: profession, gender, religion, and race.”From this paper · §Bias & Toxicity Evaluations
Show 10 more
- Megatron-LM2019 · cited 2×, 1 in Method“We trained OPT-175B on 992 80GB A100 GPUs, by utilizing Fully Sharded Data Parallel Artetxe et al. 2021 with Megatron-LM Tensor Parallelism Shoeybi et al. 2019.”From this paper · §Method
- BookCorpus (books & movies)2015 · cited 1×, 1 in Method“We included the BookCorpus Zhu et al. 2015 and Stories Trinh and Le 2018 subsets of the RoBERTa corpus and utilized an updated version of CCNews, containing news stories crawled through September 28, 2021.”From this paper · §Method
- PaLM2022 · cited 10דFollowing PaLM Chowdhery et al. 2022, we sample 25 generations of 20 tokens using nucleus sampling Holtzman et al. 2020 (p=0.9p=0.9) for each of 10,00010,000 randomly sampled prompts from RTP, and report mean toxicity pr…”From this paper · §Bias & Toxicity Evaluations
- Gopher2021 · cited 9דLarge language models (LLMs) trained on massive text collections have shown surprising emergent capabilities to generate text and perform zero- and few-shot learning Brown et al. 2020; Lieber et al. 2021; Smith et al. 20…”From this paper · §Introduction
- LaMDA2022 · cited 6דSimilar to other LLMs, OPT-175B can produce factually incorrect statements Adiwardana et al. 2020; Brown et al. 2020; Roller et al. 2021; Rae et al. 2021; Chowdhery et al. 2022; Thoppilan et al. 2022.”From this paper · §Limitations
- Megatron-Turing NLG2022 · cited 4דAuto-regressive language models Mikolov et al. 2009 have seen the largest growth in model size, from 117M parameters Radford et al. 2018 to over 500B parameters Smith et al. 2022; Chowdhery et al. 2022.”From this paper · §Related Work
- Chinchilla2022 · cited 4דThese scaling gains come not only from growing the total number of parameters in the models, but also the amount and quality of pre-training data Liu et al. 2019b; Hoffmann et al. 2022.”From this paper · §Related Work
- GPT-NeoX-20B2022 · cited 4דThere are a few notable efforts towards open sourcing LLMs from non-profit research organizations including EleutherAI Black et al. 2022 and BigScience.1111 11 https://huggingface.co/bigscience/tr11-176B-ml-logs/tensorbo…”From this paper · §Related Work
- InstructGPT2022 · cited 3דRecent efforts have shown gains by fine-tuning models to directly respond to instruction-style prompting Wei et al. 2021; Min et al. 2021; Sanh et al. 2021; Ouyang et al. 2022.”From this paper · §Related Work
- FLAN2021 · cited 2דRecent efforts have shown gains by fine-tuning models to directly respond to instruction-style prompting Wei et al. 2021; Min et al. 2021; Sanh et al. 2021; Ouyang et al. 2022.”From this paper · §Related Work
Led to
- BLIP-22023 · cited 3×, 1 in Method“For the frozen language model, we explore the unsupervised-trained OPT model family (Zhang et al. 2022) for decoder-based LLMs, and the instruction-trained FlanT5 model family (Chung et al. 2022) for encoder-decoder-base…”From BLIP-2 · §Method
Abstract
Large language models, which are often trained for hundreds of thousands of compute days, have shown remarkable capabilities for zero- and few-shot learning. Given their computational cost, these models are difficult to replicate without significant capital. For the few that are available through APIs, no access is granted to the full model weights, making them difficult to study. We present Open Pre-trained Transformers (OPT), a suite of decoder-only pre-trained transformers ranging from 125M to 175B parameters, which we aim to fully and responsibly share with interested researchers. We show that OPT-175B is comparable to GPT-3, while requiring only 1/7th the carbon footprint to develop. We are also releasing our logbook detailing the infrastructure challenges we faced, along with code for experimenting with all of the released models.