Paper Lineage
Esc
BenchmarkNov 2018arXiv 1811.00491cs.CL

A Corpus for Reasoning About Natural Language Grounded in Photographs

Alane Suhr, Stephanie Zhou, Ally Zhang and 3 others

We introduce a new dataset for joint reasoning about natural language and images, with a focus on semantic diversity, compositionality, and visual reasoning challenges. The data contains 107,292 examples of English sentences paired with web photographs.

From the abstract

Built on

3 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (3)

  • VQA2015 · cited 3×
    “However, commonly used resources for language and vision (Antol et al. 2015; Chen et al. 2016, e.g.,) focus mostly on identification of object properties and few spatial relations (Ferraro et al. 2015; Alikhani and Stone…”
    From this paper · §Introduction
  • MS COCO2014 · cited 2×
    “We detect the objects in the images using a Mask R-CNN model He et al. 2017; Girshick et al. 2018 pre-trained on the COCO detection task Lin et al. 2014.”
    From this paper · §Evaluation Systems
  • VQA v22016 · cited 2×
    “This relatively simple reasoning, together with biases in the data, removes much of the need to consider language compositionality Goyal et al. 2017.”
    From this paper · §Introduction

Led to

  • VisualBERT2019 · cited 4×
    “We evaluate VisualBERT on four different types of vision-and-language applications: (1) Visual Question Answering (VQA 2.0) (Goyal et al. 2017), (2) Visual Commonsense Reasoning (VCR) (Zellers et al. 2019), (3) Natural L…”
    From VisualBERT · §Experiment
  • LXMERT2019 · cited 3×, 2 in Method
    “NLVR2NLVR^{2} Suhr et al. 2019 is a challenging visual reasoning dataset where some existing approaches Hu et al. 2017; Perez et al. 2018 fail, and the SotA method is ‘MaxEnt’ in Suhr et al. 2019.”
    From LXMERT · §Experimental Setup and Results
  • VL-BERT meta-analysis2020 · cited 1×, 1 in Method
    “We consider the most common tasks used to evaluate V&L BERTs, spanning four groups: vocab-based VQA (Goyal et al. 2017; Hudson and Manning 2019), image–text retrieval (Lin et al. 2014; Plummer et al. 2015), referring exp…”
    From VL-BERT meta-analysis · §Experimental Setup
  • VinVL2021 · cited 2×
    “We then fine-tune the pre-trained Oscar+ for a wide range of downstream tasks, including VL understanding tasks such as VQA [8], GQA [13], NLVR2 [35], and COCO text-image retrieval [25], and VL generation tasks such as C…”
    From VinVL · §Introduction
  • VL-T52021 · cited 3×, 1 in Method
    “Image ids are used to discriminate regions from different images, and is used when multiple images are given to the model (i.e., in NLVR2NLVR^{2} (Suhr et al. 2019), models take two input images).”
    From VL-T5 · §Model
  • ViLT2021 · cited 3×, 2 in Method
    “We evaluate ViLT on two widely explored types of vision-and-language downstream tasks: for classification, we use VQAv2 (Goyal et al. 2017) and NLVR2 (Suhr et al. 2018), and for retrieval, we use MSCOCO and Flickr30K (F3…”
    From ViLT · §Experiments
  • ALBEF2021 · cited 2×
    “NLVR2 [19], VQA [20]), but most of them require high-resolution input images and pre-trained object detectors.”
    From ALBEF · §Related Work
  • VLMo2021 · cited 2×
    “We first conduct fine-tuning experiments on two widely used classification datasets: visual question answering [15] and natural language for visual reasoning [42].”
    From VLMo · §Experiments
  • METER2021 · cited 2×, 1 in Method
    “We test them on visual question answering antol2015vqa, visual reasoning suhr2018corpus, image-text retrieval lin2014microsoft; plummer2015flickr30k, and visual entailment xie2019visual tasks.”
    From METER · §Introduction
  • VL-BEiT2022 · cited 2×
    “We use NLVR2 [36] dataset to evaluate the model.”
    From VL-BEiT · §Experiments
  • BEiT-32022 · cited 2×
    “We evaluate the capabilities of BEiT-3 on the widely used vision-language understanding and generation benchmarks, including visual question answering [19], visual reasoning [49], image-text retrieval [41, 31], and image…”
    From BEiT-3 · §Experiments on Vision and Vision-Language Tasks
Abstract

We introduce a new dataset for joint reasoning about natural language and images, with a focus on semantic diversity, compositionality, and visual reasoning challenges. The data contains 107,292 examples of English sentences paired with web photographs. The task is to determine whether a natural language caption is true about a pair of photographs. We crowdsource the data using sets of visually rich images and a compare-and-contrast task to elicit linguistically diverse language. Qualitative analysis shows the data requires compositional joint reasoning, including about quantities, comparisons, and relations. Evaluation using state-of-the-art visual reasoning methods shows the data presents a strong challenge.