How Can We Know What Language Models Know?
Recent work has presented intriguing results examining the knowledge contained in language models (LM) by having the LM fill in the blanks of prompts such as "Obama is a _ by profession". These prompts are usually manually created, and quite possibly sub-optimal; another prompt such as "Obama worked as a _" may result in more accurately predicting the correct profession.
Also cited · not yet reviewed (2)
- BERT2018 · cited 4דAs for the models to probe, in our main experiments we use the standard BERT-base and BERT-large models (Devlin et al. 2019).”From this paper · §Main Experiments
- decaNLP2018 · cited 2דIn previous work (McCann et al. 2018; Radford et al. 2019; Petroni et al. 2019), trt_{r} has been a single manually defined prompt based on the intuition of the experimenter.”From this paper · §Knowledge Retrieval from LMs
Led to
- Knowledge in LM parameters2020 · cited 2דPast work investigating “language models as knowledge bases” has typically tried to understand the scope of the information stored in the model using synthetic tasks that are similar to the pre-training objective Petroni…”From Knowledge in LM parameters · §Introduction
- True few-shot learning2021 · cited 4דHowever, the few-shot performance of LMs is very sensitive to the textual task description [3, 4, 5, 6, “prompt”;], order of training examples [6, 7, 8], decoding strategy [9, 10], and other hyperparameters [3, 5, 9, 11,…”From True few-shot learning · §Introduction
- BIG-bench2022 · cited 2×, 2 in Method“Furthermore, Jiang et al. 2020 studied calibration on generative language models (T5, BART, and GPT-2) and found that their predictive probabilities on question-answering tasks are not well calibrated.”From BIG-bench · §Behavior of language models and human raters on BIG-bench
Abstract
Recent work has presented intriguing results examining the knowledge contained in language models (LM) by having the LM fill in the blanks of prompts such as "Obama is a _ by profession". These prompts are usually manually created, and quite possibly sub-optimal; another prompt such as "Obama worked as a _" may result in more accurately predicting the correct profession. Because of this, given an inappropriate prompt, we might fail to retrieve facts that the LM does know, and thus any given prompt only provides a lower bound estimate of the knowledge contained in an LM. In this paper, we attempt to more accurately estimate the knowledge contained in LMs by automatically discovering better prompts to use in this querying process. Specifically, we propose mining-based and paraphrasing-based methods to automatically generate high-quality and diverse prompts, as well as ensemble methods to combine answers from different prompts. Extensive experiments on the LAMA benchmark for extracting relational knowledge from LMs demonstrate that our methods can improve accuracy from 31.1% to 39.6%, providing a tighter lower bound on what LMs know. We have released the code and the resulting LM Prompt And Query Archive (LPAQA) at https://github.com/jzbjyb/LPAQA.