Paper Lineage
Esc
BenchmarkOct 2018arXiv 1810.12885cs.CL

ReCoRD: Bridging the Gap between Human and Machine Commonsense Reading Comprehension

Sheng Zhang, Xiaodong Liu, Jingjing Liu and 3 others

We present a large-scale dataset, ReCoRD, for machine reading comprehension requiring commonsense reasoning. Experiments on this dataset demonstrate that the performance of state-of-the-art MRC systems fall far behind human performance.

From the abstract

Built on

2 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (2)

  • SQuAD2016 · cited 5×
    “Since all candidate answers were extracted from in the passage, ReCoRD can also be formalized as a extractive MRC dataset, similar to SQuAD Rajpurkar et al. 2016 and NewsQA Trischler et al. 2017.”
    From this paper · §Related Datasets
  • ELMo2018 · cited 2×
    “We also evaluate DocQA with ELMo Peters et al. 2018 to analyze the impact of largely pre-trained encoder on our dataset.”
    From this paper · §Evaluation

Led to

  • T52019 · cited 1×, 1 in Method
    “Question answering (MultiRC (Khashabi et al. 2018), ReCoRD (Zhang et al. 2018), BoolQ (Clark et al. 2019))”
    From T5 · §Setup
  • Knowledge in LM parameters2020 · cited 3×, 2 in Method
    “Most past work on question answering either explicitly feeds pertinent information to the model alongside the question (for example, an article that contains the answer Rajpurkar et al. 2016; Zhang et al. 2018; Khashabi…”
    From Knowledge in LM parameters · §Introduction
  • GLaM2021 · cited 1×, 1 in Method
    “On a few tasks, such as ReCoRD (Zhang et al. 2018) and COPA (Gordon et al. 2012), the non-normalized loss can yield better results and thus is adopted.”
    From GLaM · §Experiment Setup
  • Fairseq MoE LMs2021 · cited 1×, 1 in Method
    “Zero-shot: in addition to the 3 few-shot tasks, we evaluate on ReCoRD (Zhang et al. 2018), HellaSwag (Zellers et al. 2019) and PIQA (Bisk et al. 2020).”
    From Fairseq MoE LMs · §Experimental Setup
Abstract

We present a large-scale dataset, ReCoRD, for machine reading comprehension requiring commonsense reasoning. Experiments on this dataset demonstrate that the performance of state-of-the-art MRC systems fall far behind human performance. ReCoRD represents a challenge for future research to bridge the gap between human and machine commonsense reading comprehension. ReCoRD is available at http://nlp.jhu.edu/record.