SQuAD: 100,000+ Questions for Machine Comprehension of Text
We present the Stanford Question Answering Dataset (SQuAD), a new reading comprehension dataset consisting of 100,000+ questions posed by crowdworkers on a set of Wikipedia articles, where the answer to each question is a segment of text from the corresponding reading passage. We analyze the dataset to understand the types of reasoning required to answer the questions, leaning heavily on dependency and constituency trees.
SQuAD has no earlier papers in this dataset.
Led to
- DrQA2017 · cited 6דThat subfield has made considerable progress recently thanks to new deep learning architectures like attention-based and memory-augmented neural networks Bahdanau et al. 2015; Weston et al. 2015; Graves et al. 2014 and r…”From DrQA · §Related Work
- CoQA2018 · cited 3×, 1 in Method“We use the Document Reader (DrQA) model of Chen et al. 2017, which has demonstrated strong performance on multiple datasets Rajpurkar et al. 2016; Labutov et al. 2018.”From CoQA · §Models
- BERT2018 · cited 3דWhen integrating contextual word embeddings with existing task-specific architectures, ELMo advances the state of the art for several major NLP benchmarks Peters et al. 2018a including question answering Rajpurkar et al.…”From BERT · §Related Work
- ReCoRD2018 · cited 5דSince all candidate answers were extracted from in the passage, ReCoRD can also be formalized as a extractive MRC dataset, similar to SQuAD Rajpurkar et al. 2016 and NewsQA Trischler et al. 2017.”From ReCoRD · §Related Datasets
- UniLM2019 · cited 3דThe task is to answer a question given a passage [33, 34, 15].”From UniLM · §Experiments
- BoolQ2019 · cited 2×, 1 in Method“First, we use the QNLI task from GLUE Wang et al. 2018, where the model must determine if a sentence from SQuAD 1.1 Rajpurkar et al. 2016 contains the answer to an input question or not.”From BoolQ · §Training Yes/No QA Models
- RoBERTa2019 · cited 2×, 2 in Method“The General Language Understanding Evaluation (GLUE) benchmark Wang et al. 2019b is a collection of 9 datasets for evaluating natural language understanding systems.66 6 The datasets are: CoLA Warstadt et al. 2018, Stanf…”From RoBERTa · §Experimental Setup
- T52019 · cited 3×, 2 in Method“Natural language inference (MNLI (Williams et al. 2017), QNLI (Rajpurkar et al. 2016), RTE (Dagan et al. 2005), CB (De Marneff et al. 2019))”From T5 · §Setup
- BART2019 · cited 2×, 1 in Method“Rajpurkar et al. 2016a an extractive question answering task on Wikipedia paragraphs.”From BART · §Comparing Pre-training Objectives
- Knowledge in LM parameters2020 · cited 4×, 2 in Method“Most past work on question answering either explicitly feeds pertinent information to the model alongside the question (for example, an article that contains the answer Rajpurkar et al. 2016; Zhang et al. 2018; Khashabi…”From Knowledge in LM parameters · §Introduction
- BIG-bench2022 · cited 8דFor instance, benchmarks often propose tasks that codify narrow subsets of areas, such as language understanding (Wang et al. 2019a), summarization (See et al. 2017; Hermann et al. 2015; Narayan et al. 2018; Koupaee & Wa…”From BIG-bench · §Introduction
Abstract
We present the Stanford Question Answering Dataset (SQuAD), a new reading comprehension dataset consisting of 100,000+ questions posed by crowdworkers on a set of Wikipedia articles, where the answer to each question is a segment of text from the corresponding reading passage. We analyze the dataset to understand the types of reasoning required to answer the questions, leaning heavily on dependency and constituency trees. We build a strong logistic regression model, which achieves an F1 score of 51.0%, a significant improvement over a simple baseline (20%). However, human performance (86.8%) is much higher, indicating that the dataset presents a good challenge problem for future research. The dataset is freely available at https://stanford-qa.com