Paper Lineage
Esc
BenchmarkMay 2019arXiv 1905.10044cs.CL

BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Christopher Clark, Kenton Lee, Ming-Wei Chang and 3 others

In this paper we study yes/no questions that are naturally occurring --- meaning that they are generated in unprompted and unconstrained settings. We build a reading comprehension dataset, BoolQ, of such questions, and show that they are unexpectedly challenging.

From the abstract

Built on

4 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (4)

  • ELMo2018 · cited 4×, 1 in Method
    “Our Recurrent +ELMo model uses the language model from Peters et al. 2018 to provide contextualized embeddings to the baseline model outlined above, as recommended by the authors.”
    From this paper · §Results
  • BERT2018 · cited 4×, 1 in Method
    “Unsupervised: It is well known that unsupervised pre-training using language-modeling objectives Peters et al. 2018; Devlin et al. 2018; Radford et al. 2018, can improve performance on many tasks.”
    From this paper · §Training Yes/No QA Models
  • SQuAD2016 · cited 2×, 1 in Method
    “First, we use the QNLI task from GLUE Wang et al. 2018, where the model must determine if a sentence from SQuAD 1.1 Rajpurkar et al. 2016 contains the answer to an input question or not.”
    From this paper · §Training Yes/No QA Models
  • CoQA2018 · cited 2×
    “Yes/No questions make up a subset of the reading comprehension datasets CoQA Reddy et al. 2018, QuAC Choi et al. 2018, and HotPotQA Yang et al. 2018, and are present in the ShARC Saeidi et al. 2018 dataset.”
    From this paper · §Related Work

Led to

  • T52019 · cited 1×, 1 in Method
    “Question answering (MultiRC (Khashabi et al. 2018), ReCoRD (Zhang et al. 2018), BoolQ (Clark et al. 2019))”
    From T5 · §Setup
  • Knowledge in LM parameters2020 · cited 3×, 2 in Method
    “Most past work on question answering either explicitly feeds pertinent information to the model alongside the question (for example, an article that contains the answer Rajpurkar et al. 2016; Zhang et al. 2018; Khashabi…”
    From Knowledge in LM parameters · §Introduction
  • Fairseq MoE LMs2021 · cited 1×, 1 in Method
    “In addition to our zero-shot tasks, we also evaluate on 3 widely-used classification tasks: BoolQ (Clark et al. 2019), MNLI (Williams et al. 2018) and SST-2 (Socher et al. 2013).”
    From Fairseq MoE LMs · §Experimental Setup
Abstract

In this paper we study yes/no questions that are naturally occurring --- meaning that they are generated in unprompted and unconstrained settings. We build a reading comprehension dataset, BoolQ, of such questions, and show that they are unexpectedly challenging. They often query for complex, non-factoid information, and require difficult entailment-like inference to solve. We also explore the effectiveness of a range of transfer learning baselines. We find that transferring from entailment data is more effective than transferring from paraphrase or extractive QA data, and that it, surprisingly, continues to be very beneficial even when starting from massive pre-trained language models such as BERT. Our best method trains BERT on MultiNLI and then re-trains it on our train set. It achieves 80.4% accuracy compared to 90% accuracy of human annotators (and 62% majority-baseline), leaving a significant gap for future work.