Paper Lineage
Esc
DatasetDec 2020arXiv 2101.00027cs.CL

The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Leo Gao, Stella Biderman, Sid Black and 9 others

Recent work has demonstrated that increased training dataset diversity improves general cross-domain knowledge and downstream generalization capability for large-scale language models. With this in mind, we present \textit{the Pile}: an 825 GiB English text corpus targeted at training large-scale language models.

From the abstract

Built on

7 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (7)

  • GPT-32020 · cited 15×, 1 in Method
    “Following Brown et al. 2020, we increase the weights of higher quality components, with certain high-quality datasets such as Wikipedia being seen up to 3 times (“epochs”) for each full epoch over the Pile.”
    From this paper · §The Pile Datasets
  • Kaplan scaling laws2020 · cited 2×, 1 in Method
    “Recent work has shown an increased focus on the empirical scaling laws of language models (Kaplan et al. 2020; Henighan et al. 2020).”
    From this paper · §Benchmarking Language Models with the Pile
  • T52019 · cited 5×
    “More recently, language models (Radford et al. 2018; Radford et al. 2019; Brown et al. 2020; Rosset 2019; Shoeybi et al. 2019) and masked language models (Devlin et al. 2019; Liu et al. 2019; Raffel et al. 2019) have bee…”
    From this paper · §Related Work
  • “More recently, language models (Radford et al. 2018; Radford et al. 2019; Brown et al. 2020; Rosset 2019; Shoeybi et al. 2019) and masked language models (Devlin et al. 2019; Liu et al. 2019; Raffel et al. 2019) have bee…”
    From this paper · §Related Work
Show 3 more
  • BERT2018 · cited 2×
    “More recently, language models (Radford et al. 2018; Radford et al. 2019; Brown et al. 2020; Rosset 2019; Shoeybi et al. 2019) and masked language models (Devlin et al. 2019; Liu et al. 2019; Raffel et al. 2019) have bee…”
    From this paper · §Related Work
  • RoBERTa2019 · cited 2×
    “More recently, language models (Radford et al. 2018; Radford et al. 2019; Brown et al. 2020; Rosset 2019; Shoeybi et al. 2019) and masked language models (Devlin et al. 2019; Liu et al. 2019; Raffel et al. 2019) have bee…”
    From this paper · §Related Work
  • Megatron-LM2019 · cited 2×
    “More recently, language models (Radford et al. 2018; Radford et al. 2019; Brown et al. 2020; Rosset 2019; Shoeybi et al. 2019) and masked language models (Devlin et al. 2019; Liu et al. 2019; Raffel et al. 2019) have bee…”
    From this paper · §Related Work

Led to

  • Codex2021 · cited 2×
    “More recently, language models have also fueled progress towards the longstanding challenge of program synthesis (Simon 1963; Manna & Waldinger 1971), spurred by the presence of code in large datasets (Husain et al. 2019…”
    From Codex · §Introduction
  • Gopher2021 · cited 1×, 1 in Method
    “Since GPT-3 there has been a 178B parameter Transformer language model Jurassic-1 (Lieber et al. 2021) which uses a diverse training set and a larger tokenizer vocabulary size, along with an announced 530B Megatron-Turin…”
    From Gopher · §Background
  • Fairseq MoE LMs2021 · cited 3×, 2 in Method
    “For out-of-domain we use data from The Pile (Gao et al. 2021), a public dataset that combines data from 22 diverse sources (e.g., ArXiv, Github, OpenSubtitles, etc.).”
    From Fairseq MoE LMs · §Experimental Setup
  • Megatron-Turing NLG2022 · cited 5×, 5 in Method
    “To compile our training dataset, we made use of recent work aimed at collecting a diverse training set for language modeling [17].”
    From Megatron-Turing NLG · §Training Dataset and Model Configuration
  • GPT-NeoX-20B2022 · cited 4×
    “GPT-NeoX-20B was trained on the Pile (Gao et al. 2020), a massive curated dataset designed specifically for training large language models.”
    From GPT-NeoX-20B · §Training
  • OPT2022 · cited 2×, 2 in Method
    “The pre-training corpus contains a concatenation of datasets used in RoBERTa Liu et al. 2019b, the Pile Gao et al. 2021a, and PushShift.io Reddit Baumgartner et al. 2020; Roller et al. 2021.”
    From OPT · §Method
Abstract

Recent work has demonstrated that increased training dataset diversity improves general cross-domain knowledge and downstream generalization capability for large-scale language models. With this in mind, we present \textit{the Pile}: an 825 GiB English text corpus targeted at training large-scale language models. The Pile is constructed from 22 diverse high-quality subsets -- both existing and newly constructed -- many of which derive from academic or professional sources. Our evaluation of the untuned performance of GPT-2 and GPT-3 on the Pile shows that these models struggle on many of its components, such as academic writing. Conversely, models trained on the Pile improve significantly over both Raw CC and CC-100 on all components of the Pile, while improving performance on downstream evaluations. Through an in-depth exploratory analysis, we document potentially concerning aspects of the data for prospective users. We make publicly available the code used in its construction.