The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Recent work has demonstrated that increased training dataset diversity improves general cross-domain knowledge and downstream generalization capability for large-scale language models. With this in mind, we present \textit{the Pile}: an 825 GiB English text corpus targeted at training large-scale language models.
Also cited · not yet reviewed (7)
- GPT-32020 · cited 15×, 1 in Method“Following Brown et al. 2020, we increase the weights of higher quality components, with certain high-quality datasets such as Wikipedia being seen up to 3 times (“epochs”) for each full epoch over the Pile.”From this paper · §The Pile Datasets
- Kaplan scaling laws2020 · cited 2×, 1 in Method“Recent work has shown an increased focus on the empirical scaling laws of language models (Kaplan et al. 2020; Henighan et al. 2020).”From this paper · §Benchmarking Language Models with the Pile
- T52019 · cited 5דMore recently, language models (Radford et al. 2018; Radford et al. 2019; Brown et al. 2020; Rosset 2019; Shoeybi et al. 2019) and masked language models (Devlin et al. 2019; Liu et al. 2019; Raffel et al. 2019) have bee…”From this paper · §Related Work
- BookCorpus (books & movies)2015 · cited 3דMore recently, language models (Radford et al. 2018; Radford et al. 2019; Brown et al. 2020; Rosset 2019; Shoeybi et al. 2019) and masked language models (Devlin et al. 2019; Liu et al. 2019; Raffel et al. 2019) have bee…”From this paper · §Related Work
Show 3 more
- BERT2018 · cited 2דMore recently, language models (Radford et al. 2018; Radford et al. 2019; Brown et al. 2020; Rosset 2019; Shoeybi et al. 2019) and masked language models (Devlin et al. 2019; Liu et al. 2019; Raffel et al. 2019) have bee…”From this paper · §Related Work
- RoBERTa2019 · cited 2דMore recently, language models (Radford et al. 2018; Radford et al. 2019; Brown et al. 2020; Rosset 2019; Shoeybi et al. 2019) and masked language models (Devlin et al. 2019; Liu et al. 2019; Raffel et al. 2019) have bee…”From this paper · §Related Work
- Megatron-LM2019 · cited 2דMore recently, language models (Radford et al. 2018; Radford et al. 2019; Brown et al. 2020; Rosset 2019; Shoeybi et al. 2019) and masked language models (Devlin et al. 2019; Liu et al. 2019; Raffel et al. 2019) have bee…”From this paper · §Related Work
Led to
- Codex2021 · cited 2דMore recently, language models have also fueled progress towards the longstanding challenge of program synthesis (Simon 1963; Manna & Waldinger 1971), spurred by the presence of code in large datasets (Husain et al. 2019…”From Codex · §Introduction
- Gopher2021 · cited 1×, 1 in Method“Since GPT-3 there has been a 178B parameter Transformer language model Jurassic-1 (Lieber et al. 2021) which uses a diverse training set and a larger tokenizer vocabulary size, along with an announced 530B Megatron-Turin…”From Gopher · §Background
- Fairseq MoE LMs2021 · cited 3×, 2 in Method“For out-of-domain we use data from The Pile (Gao et al. 2021), a public dataset that combines data from 22 diverse sources (e.g., ArXiv, Github, OpenSubtitles, etc.).”From Fairseq MoE LMs · §Experimental Setup
- Megatron-Turing NLG2022 · cited 5×, 5 in Method“To compile our training dataset, we made use of recent work aimed at collecting a diverse training set for language modeling [17].”From Megatron-Turing NLG · §Training Dataset and Model Configuration
- GPT-NeoX-20B2022 · cited 4דGPT-NeoX-20B was trained on the Pile (Gao et al. 2020), a massive curated dataset designed specifically for training large language models.”From GPT-NeoX-20B · §Training
- OPT2022 · cited 2×, 2 in Method“The pre-training corpus contains a concatenation of datasets used in RoBERTa Liu et al. 2019b, the Pile Gao et al. 2021a, and PushShift.io Reddit Baumgartner et al. 2020; Roller et al. 2021.”From OPT · §Method
Abstract
Recent work has demonstrated that increased training dataset diversity improves general cross-domain knowledge and downstream generalization capability for large-scale language models. With this in mind, we present \textit{the Pile}: an 825 GiB English text corpus targeted at training large-scale language models. The Pile is constructed from 22 diverse high-quality subsets -- both existing and newly constructed -- many of which derive from academic or professional sources. Our evaluation of the untuned performance of GPT-2 and GPT-3 on the Pile shows that these models struggle on many of its components, such as academic writing. Conversely, models trained on the Pile improve significantly over both Raw CC and CC-100 on all components of the Pile, while improving performance on downstream evaluations. Through an in-depth exploratory analysis, we document potentially concerning aspects of the data for prospective users. We make publicly available the code used in its construction.