GPT-NeoX-20B: An Open-Source Autoregressive Language Model
We introduce GPT-NeoX-20B, a 20 billion parameter autoregressive language model trained on the Pile, whose weights will be made freely and openly available to the public through a permissive license. It is, to the best of our knowledge, the largest dense autoregressive model that has publicly available weights at the time of submission.
Also cited · not yet reviewed (10)
- GPT-32020 · cited 10×, 2 in Method“GPT-NeoX-20B is an autoregressive transformer decoder model whose architecture largely follows that of GPT-3 (Brown et al. 2020), with a few notable deviations described below.”From this paper · §Model Design and Implementation
- Megatron-LM2019 · cited 3×, 2 in Method“Our model is trained using a codebase that builds on Megatron (Shoeybi et al. 2020) and DeepSpeed (Rasley et al. 2020) to facilitate efficient and straightforward training of large language models with tens of billions o…”From this paper · §Model Design and Implementation
- RoFormer (RoPE)2021 · cited 2×, 2 in Method“We use rotary embeddings (Su et al. 2021) instead of the learned positional embeddings used in GPT models (Radford et al. 2018), based on our positive prior experiences using it in training LLMs.”From this paper · §Model Design and Implementation
- Kaplan scaling laws2020 · cited 3×, 1 in Method“Our model has 20 billion parameters, of which 19.9 billion are “non-embedding” parameters that Kaplan et al. 2020 identify as the proper number to use for scaling laws analysis.”From this paper · §Model Design and Implementation
Show 6 more
- The Pile2020 · cited 4דGPT-NeoX-20B was trained on the Pile (Gao et al. 2020), a massive curated dataset designed specifically for training large language models.”From this paper · §Training
- MMLU2020 · cited 3דTo do this, we use a dataset of multiple choice questions in a variety of diverse domains developed by Hendrycks et al. 2021a.”From this paper · §Performance Evaluations
- T52019 · cited 2דOver the past several years, there has been an explosion in research surrounding large language models (LLMs) for natural language processing, catalyzed largely by the impressive performance of Transformer-based language…”From this paper · §Introduction
- Fairseq MoE LMs2021 · cited 2דWe compare with the GPT-3 models on the OpenAI API (Brown et al. 2020), the open source FairSeq dense models (Artetxe et al. 2021), and GPT-J-6B (Wang and Komatsuzaki 2021).”From this paper · §Performance Evaluations
- Megatron-Turing NLG2022 · cited 2דA consequence of this has been an abundance of research focusing on scaling Transformer models up to ever-larger scales, resulting in dense models that surpass 500B parameters (Smith et al. 2022; Chowdhery et al. 2022),…”From this paper · §Introduction
- PaLM2022 · cited 2דA consequence of this has been an abundance of research focusing on scaling Transformer models up to ever-larger scales, resulting in dense models that surpass 500B parameters (Smith et al. 2022; Chowdhery et al. 2022),…”From this paper · §Introduction
Led to
- OPT2022 · cited 4דThere are a few notable efforts towards open sourcing LLMs from non-profit research organizations including EleutherAI Black et al. 2022 and BigScience.1111 11 https://huggingface.co/bigscience/tr11-176B-ml-logs/tensorbo…”From OPT · §Related Work
- BIG-bench2022 · cited 2דThe quantitative and qualitative changes that occur in language models as they become larger are potentially transformative (Bommasani et al. 2021; Black et al. 2022).”From BIG-bench · §Introduction
Abstract
We introduce GPT-NeoX-20B, a 20 billion parameter autoregressive language model trained on the Pile, whose weights will be made freely and openly available to the public through a permissive license. It is, to the best of our knowledge, the largest dense autoregressive model that has publicly available weights at the time of submission. In this work, we describe \model{}'s architecture and training and evaluate its performance on a range of language-understanding, mathematics, and knowledge-based tasks. We find that GPT-NeoX-20B is a particularly powerful few-shot reasoner and gains far more in performance when evaluated five-shot than similarly sized GPT-3 and FairSeq models. We open-source the training and evaluation code, as well as the model weights, at https://github.com/EleutherAI/gpt-neox.