How Much Knowledge Can You Pack Into the Parameters of a Language Model?
It has recently been observed that neural language models trained on unstructured text can implicitly store and retrieve knowledge using natural language queries. In this short paper, we measure the practical utility of this approach by fine-tuning pre-trained models to answer questions without access to any external context or knowledge.
Also cited · not yet reviewed (11)
- T52019 · cited 9×, 3 in Method“Big, deep neural language models that have been pre-trained on unlabeled text have proven to be extremely performant when fine-tuned on downstream Natural Language Processing (NLP) tasks Devlin et al. 2018; Yang et al. 2…”From this paper · §Introduction
- SQuAD2016 · cited 4×, 2 in Method“Most past work on question answering either explicitly feeds pertinent information to the model alongside the question (for example, an article that contains the answer Rajpurkar et al. 2016; Zhang et al. 2018; Khashabi…”From this paper · §Introduction
- BERT2018 · cited 4×, 2 in Method“Big, deep neural language models that have been pre-trained on unlabeled text have proven to be extremely performant when fine-tuned on downstream Natural Language Processing (NLP) tasks Devlin et al. 2018; Yang et al. 2…”From this paper · §Introduction
- ReCoRD2018 · cited 3×, 2 in Method“Most past work on question answering either explicitly feeds pertinent information to the model alongside the question (for example, an article that contains the answer Rajpurkar et al. 2016; Zhang et al. 2018; Khashabi…”From this paper · §Introduction
Show 7 more
- BoolQ2019 · cited 3×, 2 in Method“Most past work on question answering either explicitly feeds pertinent information to the model alongside the question (for example, an article that contains the answer Rajpurkar et al. 2016; Zhang et al. 2018; Khashabi…”From this paper · §Introduction
- XLNet2019 · cited 2×, 1 in Method“Big, deep neural language models that have been pre-trained on unlabeled text have proven to be extremely performant when fine-tuned on downstream Natural Language Processing (NLP) tasks Devlin et al. 2018; Yang et al. 2…”From this paper · §Introduction
- ALBERT2019 · cited 2×, 1 in Method“Big, deep neural language models that have been pre-trained on unlabeled text have proven to be extremely performant when fine-tuned on downstream Natural Language Processing (NLP) tasks Devlin et al. 2018; Yang et al. 2…”From this paper · §Introduction
- Transformer2017 · cited 1×, 1 in Method“Currently, the most popular model architectures used in transfer learning for NLP are Transformer-based Vaswani et al. 2017 “encoder-only” models like BERT Devlin et al. 2018.”From this paper · §Background
- ELMo2018 · cited 1×, 1 in Method“The popularity of this form of “transfer learning” is attributable to its empirical success on many NLP tasks Peters et al. 2018; Devlin et al. 2018; Yang et al. 2019; Lan et al. 2019; Raffel et al. 2019.”From this paper · §Background
- RoBERTa2019 · cited 2דBig, deep neural language models that have been pre-trained on unlabeled text have proven to be extremely performant when fine-tuned on downstream Natural Language Processing (NLP) tasks Devlin et al. 2018; Yang et al. 2…”From this paper · §Introduction
- LPAQA (What LMs know)2019 · cited 2דPast work investigating “language models as knowledge bases” has typically tried to understand the scope of the information stored in the model using synthetic tasks that are similar to the pre-training objective Petroni…”From this paper · §Introduction
Led to
- GPT-32020 · cited 4דRecent efforts include [116, 115], which fine-tuned an 11 billion parameter language model, and [33], which focused on attending over a large corpus of data at test time.”From GPT-3 · §Related Work
- Switch Transformer2021 · cited 2דAnd as in Roberts et al. 2020, we evaluate the knowledge of our models by fine-tuning on three closed-book question answering data sets: Natural Questions (Kwiatkowski et al. 2019), Web Questions (Berant et al. 2013) and…”From Switch Transformer · §Downstream Results
- Frozen2021 · cited 3×, 1 in Method“We use a 7 billion parameter transformer trained on the public dataset C4 [30] – previous work has shown that the multi-billion parameter scale is sufficient to exhibit the key capacities we are interested in studying [2…”From Frozen · §The Frozen Method
Abstract
It has recently been observed that neural language models trained on unstructured text can implicitly store and retrieve knowledge using natural language queries. In this short paper, we measure the practical utility of this approach by fine-tuning pre-trained models to answer questions without access to any external context or knowledge. We show that this approach scales with model size and performs competitively with open-domain systems that explicitly retrieve answers from an external knowledge source when answering questions. To facilitate reproducibility and future work, we release our code and trained models at https://goo.gle/t5-cbqa.