Paper Lineage
Esc
DatasetFeb 2016arXiv 1602.07332cs.CV

Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations

Ranjay Krishna, Yuke Zhu, Oliver Groth and 9 others

Despite progress in perceptual tasks such as image classification, computers still perform poorly on cognitive tasks such as image description and question answering. Cognition is core to tasks that involve not just recognizing, but reasoning about our visual world.

From the abstract

Built on

12 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (12)

  • VQA2015 · cited 16×
    “In addition, inspired by VQA Antol et al., 2015, we also collect an average of 1717 question-answer pairs based on the descriptions for each image.”
    From this paper · §Introduction
  • MS COCO2014 · cited 15×
    “MS-COCO Lin et al., 2014 recently released its dataset, with over 328,000328,000 images with sentence descriptions and segmentations of 9191 object categories.”
    From this paper · §Related Work
  • Multi-World QA2014 · cited 9×
    “Visual question answering (QA) has been recently proposed as a proxy task of evaluating a computer vision system’s ability to understand an image beyond object recognition Geman et al., 2015; Malinowski and Fritz, 2014.”
    From this paper · §Related Work
  • “We train one of the top 16 state-of-the-art image caption generator Karpathy and Fei-Fei, 2014 on (1) our dataset to generate region descriptions and on (2) Flickr30K Young et al., 2014 to generate sentence descriptions.”
    From this paper · §Experiments
Show 8 more
  • FM-IQA2015 · cited 8×
    “Most new datasets Yu et al., 2015; Ren et al., 2015a; Antol et al., 2015; Gao et al., 2015 have collected QA pairs on MS-COCO images, either generated automatically by NLP tools Ren et al., 2015a or written by human work…”
    From this paper · §Related Work
  • Image QA models & data2015 · cited 6×
    “Most new datasets Yu et al., 2015; Ren et al., 2015a; Antol et al., 2015; Gao et al., 2015 have collected QA pairs on MS-COCO images, either generated automatically by NLP tools Ren et al., 2015a or written by human work…”
    From this paper · §Related Work
  • VGG2014 · cited 5×
    “Much progress has been made in recent years towards this goal, including image classification Deng et al., 2009; Perronnin et al., 2010; Simonyan and Zisserman, 2014; Krizhevsky et al., 2012; Szegedy et al., 2014 and obj…”
    From this paper · §Introduction
  • YFCC100M2015 · cited 4×
    “YFCC100M Thomee et al., 2016 is another large database of 100100 million images that is still largely unexplored.”
    From this paper · §Related Work
  • COCO Captions2015 · cited 4×
    “This task is closely related to image captioning Chen et al., 2015; however, results from the two are not directly comparable, as region descriptions are short, incomplete sentences.”
    From this paper · §Experiments
  • OverFeat2013 · cited 2×
    “Much progress has been made in recent years towards this goal, including image classification Deng et al., 2009; Perronnin et al., 2010; Simonyan and Zisserman, 2014; Krizhevsky et al., 2012; Szegedy et al., 2014 and obj…”
    From this paper · §Introduction
  • Ask Your Neurons2015 · cited 2×
    “The proposed models range from SVM classifiers Antol et al., 2015 and probabilistic inference Malinowski and Fritz, 2014 to recurrent neural networks Gao et al., 2015; Malinowski et al., 2015; Ren et al., 2015a and convo…”
    From this paper · §Related Work
  • Faster R-CNN2015 · cited 2×
    “Much progress has been made in recent years towards this goal, including image classification Deng et al., 2009; Perronnin et al., 2010; Simonyan and Zisserman, 2014; Krizhevsky et al., 2012; Szegedy et al., 2014 and obj…”
    From this paper · §Introduction

Led to

  • Bottom-Up Top-Down attention2017 · cited 3×, 2 in Method
    “We then train on Visual Genome krishnavisualgenome data.”
    From Bottom-Up Top-Down attention · §Approach
  • “These features are the output of Faster R-CNN [25], pre-trained using Visual Genome [17].”
    From Bilinear Attention Networks · §Experiments
  • ViLBERT2019 · cited 2×
    “This pretrain-then-transfer learning approach to vision-and-language tasks follows naturally from its widespread use in both computer vision and natural language processing where it has become the de facto standard due t…”
    From ViLBERT · §Introduction
  • LXMERT2019 · cited 2×, 1 in Method
    “As shown in Table. 1, we aggregate pre-training data from five vision-and-language datasets whose images come from MS COCO Lin et al. 2014 or Visual Genome Krishna et al. 2017.”
    From LXMERT · §Pre-Training Strategies
  • VL-BERT2019 · cited 2×
    “Visual content embedding is produced by Faster R-CNN + ResNet-101, initialized from parameters pre-trained on Visual Genome (Krishna et al. 2017) for object detection (see BUTD (Anderson et al. 2018)).”
    From VL-BERT · §Experiment
  • Decoupled box proposals captioning2019 · cited 2×, 1 in Method
    “We reimplement the Faster R-CNN model, training it to predict both 1,600 object and 400 attribute labels in Visual Genome Krishna et al. 2017, following the standard setting from Anderson et al. 2018.”
    From Decoupled box proposals captioning · §Features and Experimental Setup
  • Grid features for VQA2020 · cited 4×
    “In fact, our ablative analysis suggests that the key factors which contributed to the high accuracy of existing bottom-up attention features are: 1) the large-scale object and attribute annotations collected in the Visua…”
    From Grid features for VQA · §Introduction
  • VILLA2020 · cited 2×
    “Implementation Details For UNITER experiments, we pre-train with the same four large-scale datasets used in the original model: COCO [36], Visual Genome (VG) [28], Conceptual Captions [58] and SBU Captions [49].”
    From VILLA · §Experiments
  • VL-BERT meta-analysis2020 · cited 1×, 1 in Method
    “V&L BERTs typically extract image features using a Faster R-CNN (Ren et al. 2015) trained on the Visual Genome dataset (VG; Krishna et al. 2017), either with a ResNet-101101 (He et al. 2016) or a ResNeXT-152152 backbone…”
    From VL-BERT meta-analysis · §Experimental Setup
  • VinVL2021 · cited 2×
    “Among the aforementioned work, a widely-used object detection (OD) model [2] is trained on the Visual Genome dataset [16].”
    From VinVL · §Introduction
  • VL-T52021 · cited 2×, 2 in Method
    “We represent an input image vv with n=36n{=}36 object regions from a Faster R-CNN (Ren et al. 2015) trained on Visual Genome (Krishna et al. 2016) for object and attribute classification (Anderson et al. 2018).”
    From VL-T5 · §Model
  • ViLT2021 · cited 2×
    “Most VLP models employ an object detector pre-trained on the Visual Genome dataset (Krishna et al. 2017) annotated with 1,600 object classes and 400 attribute classes as in Anderson et al. 2018.”
    From ViLT · §Introduction
  • Conceptual 12M2021 · cited 5×, 1 in Method
    “We train a Faster-RCNN [68] on Visual Genome [43], with a ResNet101 [29] backbone trained on JFT [30] and fine-tuned on ImageNet [69].”
    From Conceptual 12M · §Evaluating Vision-and-Language Pre-Training Data
  • CLIP2021 · cited 1×, 1 in Method
    “Existing work has mainly used three datasets, MS-COCO (Lin et al. 2014), Visual Genome (Krishna et al. 2017), and YFCC100M (Thomee et al. 2016).”
    From CLIP · §Approach
  • ALBEF2021 · cited 1×, 1 in Method
    “Following UNITER [2], we construct our pre-training data using two web datasets (Conceptual Captions [4], SBU Captions [5]) and two in-domain datasets (COCO [41] and Visual Genome [42]).”
    From ALBEF · §ALBEF Pre-training
  • METER2021 · cited 2×, 1 in Method
    “We perform the investigation by pre-training models under Meter on four commonly used image-caption datasets: COCO lin2014microsoft, Conceptual Captions sharma2018conceptual, SBU Captions ordonez2011im2text, and Visual G…”
    From METER · §Introduction
  • Florence2021 · cited 2×
    “We evaluate fine-tuning on three popular object detection datasets: COCO (Lin et al. 2015), Object365 (Shao et al. 2019), and Visual Genome (Krishna et al. 2016).”
    From Florence · §Experiments
  • ViTCAP2021 · cited 4×, 1 in Method
    “To address the issue, one can simply retrieve the concepts from the open-form captions (e.g., by extracting nouns or adjective words as keywords) as the pseudo ground-truth concepts, or alternatively leverage a pre-train…”
    From ViTCAP · §ViTCAP
  • BEiT-32022 · cited 1×, 1 in Method
    “For multimodal data, there are about 1515M images and 2121M image-text pairs collected from five public datasets: Conceptual 12M (CC12M) [11], Conceptual Captions (CC3M) [46], SBU Captions (SBU) [39], COCO [31] and Visua…”
    From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
  • BLIP-22023 · cited 1×, 1 in Method
    “We use the same pre-training dataset as BLIP with 129M images in total, including COCO (Lin et al. 2014), Visual Genome (Krishna et al. 2017), CC3M (Sharma et al. 2018), CC12M (Changpinyo et al. 2021), SBU (Ordonez et al…”
    From BLIP-2 · §Method
Abstract

Despite progress in perceptual tasks such as image classification, computers still perform poorly on cognitive tasks such as image description and question answering. Cognition is core to tasks that involve not just recognizing, but reasoning about our visual world. However, models used to tackle the rich content in images for cognitive tasks are still being trained using the same datasets designed for perceptual tasks. To achieve success at cognitive tasks, models need to understand the interactions and relationships between objects in an image. When asked "What vehicle is the person riding?", computers will need to identify the objects in an image as well as the relationships riding(man, carriage) and pulling(horse, carriage) in order to answer correctly that "the person is riding a horse-drawn carriage". In this paper, we present the Visual Genome dataset to enable the modeling of such relationships. We collect dense annotations of objects, attributes, and relationships within each image to learn these models. Specifically, our dataset contains over 100K images where each image has an average of 21 objects, 18 attributes, and 18 pairwise relationships between objects. We canonicalize the objects, attributes, relationships, and noun phrases in region descriptions and questions answer pairs to WordNet synsets. Together, these annotations represent the densest and largest dataset of image descriptions, objects, attributes, relationships, and question answers.