A Corpus for Reasoning About Natural Language Grounded in Photographs
We introduce a new dataset for joint reasoning about natural language and images, with a focus on semantic diversity, compositionality, and visual reasoning challenges. The data contains 107,292 examples of English sentences paired with web photographs.
Also cited · not yet reviewed (3)
- VQA2015 · cited 3דHowever, commonly used resources for language and vision (Antol et al. 2015; Chen et al. 2016, e.g.,) focus mostly on identification of object properties and few spatial relations (Ferraro et al. 2015; Alikhani and Stone…”From this paper · §Introduction
- MS COCO2014 · cited 2דWe detect the objects in the images using a Mask R-CNN model He et al. 2017; Girshick et al. 2018 pre-trained on the COCO detection task Lin et al. 2014.”From this paper · §Evaluation Systems
- VQA v22016 · cited 2דThis relatively simple reasoning, together with biases in the data, removes much of the need to consider language compositionality Goyal et al. 2017.”From this paper · §Introduction
Led to
- VisualBERT2019 · cited 4דWe evaluate VisualBERT on four different types of vision-and-language applications: (1) Visual Question Answering (VQA 2.0) (Goyal et al. 2017), (2) Visual Commonsense Reasoning (VCR) (Zellers et al. 2019), (3) Natural L…”From VisualBERT · §Experiment
- LXMERT2019 · cited 3×, 2 in Method“NLVR2NLVR^{2} Suhr et al. 2019 is a challenging visual reasoning dataset where some existing approaches Hu et al. 2017; Perez et al. 2018 fail, and the SotA method is ‘MaxEnt’ in Suhr et al. 2019.”From LXMERT · §Experimental Setup and Results
- VL-BERT meta-analysis2020 · cited 1×, 1 in Method“We consider the most common tasks used to evaluate V&L BERTs, spanning four groups: vocab-based VQA (Goyal et al. 2017; Hudson and Manning 2019), image–text retrieval (Lin et al. 2014; Plummer et al. 2015), referring exp…”From VL-BERT meta-analysis · §Experimental Setup
- VinVL2021 · cited 2דWe then fine-tune the pre-trained Oscar+ for a wide range of downstream tasks, including VL understanding tasks such as VQA [8], GQA [13], NLVR2 [35], and COCO text-image retrieval [25], and VL generation tasks such as C…”From VinVL · §Introduction
- VL-T52021 · cited 3×, 1 in Method“Image ids are used to discriminate regions from different images, and is used when multiple images are given to the model (i.e., in NLVR2NLVR^{2} (Suhr et al. 2019), models take two input images).”From VL-T5 · §Model
- ViLT2021 · cited 3×, 2 in Method“We evaluate ViLT on two widely explored types of vision-and-language downstream tasks: for classification, we use VQAv2 (Goyal et al. 2017) and NLVR2 (Suhr et al. 2018), and for retrieval, we use MSCOCO and Flickr30K (F3…”From ViLT · §Experiments
- ALBEF2021 · cited 2דNLVR2 [19], VQA [20]), but most of them require high-resolution input images and pre-trained object detectors.”From ALBEF · §Related Work
- VLMo2021 · cited 2דWe first conduct fine-tuning experiments on two widely used classification datasets: visual question answering [15] and natural language for visual reasoning [42].”From VLMo · §Experiments
- METER2021 · cited 2×, 1 in Method“We test them on visual question answering antol2015vqa, visual reasoning suhr2018corpus, image-text retrieval lin2014microsoft; plummer2015flickr30k, and visual entailment xie2019visual tasks.”From METER · §Introduction
- VL-BEiT2022 · cited 2דWe use NLVR2 [36] dataset to evaluate the model.”From VL-BEiT · §Experiments
- BEiT-32022 · cited 2דWe evaluate the capabilities of BEiT-3 on the widely used vision-language understanding and generation benchmarks, including visual question answering [19], visual reasoning [49], image-text retrieval [41, 31], and image…”From BEiT-3 · §Experiments on Vision and Vision-Language Tasks
Abstract
We introduce a new dataset for joint reasoning about natural language and images, with a focus on semantic diversity, compositionality, and visual reasoning challenges. The data contains 107,292 examples of English sentences paired with web photographs. The task is to determine whether a natural language caption is true about a pair of photographs. We crowdsource the data using sets of visually rich images and a compare-and-contrast task to elicit linguistically diverse language. Qualitative analysis shows the data requires compositional joint reasoning, including about quantities, comparisons, and relations. Evaluation using state-of-the-art visual reasoning methods shows the data presents a strong challenge.