Paper Lineage
Esc
DatasetMay 2015arXiv 1505.04870cs.CV

Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Bryan A. Plummer, Liwei Wang, Chris M. Cervantes and 3 others

The Flickr30k dataset has become a standard benchmark for sentence-based image description. This paper presents Flickr30k Entities, which augments the 158k captions from Flickr30k with 244k coreference chains, linking mentions of the same entities across different captions for the same image, and associating them with 276k manually annotated bounding boxes.

From the abstract

Built on

4 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (4)

  • “We build on the Flickr30k dataset (Young et al., 2014), a popular benchmark for caption generation and retrieval that has been used, among others, by Chen and Zitnick, 2015; Donahue et al., 2015; Fang et al., 2015; Gong…”
    From this paper · §Introduction
  • “We build on the Flickr30k dataset (Young et al., 2014), a popular benchmark for caption generation and retrieval that has been used, among others, by Chen and Zitnick, 2015; Donahue et al., 2015; Fang et al., 2015; Gong…”
    From this paper · §Introduction
  • Deep Fragment Embeddings2014 · cited 4×
    “We build on the Flickr30k dataset (Young et al., 2014), a popular benchmark for caption generation and retrieval that has been used, among others, by Chen and Zitnick, 2015; Donahue et al., 2015; Fang et al., 2015; Gong…”
    From this paper · §Introduction
  • VGG2014 · cited 2×
    “By using better region features (Fast RCNN (Girshick, 2015) instead of ImageNet-trained VGG (Simonyan and Zisserman, 2014)) in combination with size and color cues, we are able to improve the Recall@1 for phrase localiza…”
    From this paper · §Introduction

Led to

  • Neural Baby Talk2018 · cited 4×, 2 in Method
    “Another related line of work is on resolving referring expressions kazemzadeh2014referitgame (or description-based object retrieval plummer2015flickr30k; hu2016modeling; hu2016natural; rohrbach2016grounding – given a des…”
    From Neural Baby Talk · §Related Work
  • “Moreover, we evaluate the visual grounding of bilinear attention map on Flickr30k Entities [23] outperforming previous methods, along with 25.37% improvement of inference speed taking advantage of the processing of multi…”
    From Bilinear Attention Networks · §Introduction
  • “Recently, studies of multi-modal tasks of vision and language have made significant progress, such as image captioning Anderson_2018_CVPR; Lu_2018_CVPR, visual question answering Nguyen_2018_CVPR; Teney_2018_CVPR, visual…”
    From Multi-task hierarchical VL · §Related Work
  • VisualBERT2019 · cited 3×
    “Various tasks such as visual question answering (Antol et al. 2015; Goyal et al. 2017), textual grounding (Kazemzadeh et al. 2014; Plummer et al. 2015), and visual reasoning (Suhr et al. 2019; Zellers et al. 2019) have b…”
    From VisualBERT · §Related Work
  • VL-BERT meta-analysis2020 · cited 1×, 1 in Method
    “We consider the most common tasks used to evaluate V&L BERTs, spanning four groups: vocab-based VQA (Goyal et al. 2017; Hudson and Manning 2019), image–text retrieval (Lin et al. 2014; Plummer et al. 2015), referring exp…”
    From VL-BERT meta-analysis · §Experimental Setup
  • ALIGN2021 · cited 2×, 2 in Method
    “Two benchmark datasets are considered: Flickr30K (Plummer et al. 2015) and MSCOCO (Chen et al. 2015).”
    From ALIGN · §Pre-training and Task Transfer
  • Conceptual 12M2021 · cited 1×, 1 in Method
    “The Flickr30K dataset [63] consists of 31,000 images from Flickr, each associated with five captions.”
    From Conceptual 12M · §Evaluating Vision-and-Language Pre-Training Data
  • METER2021 · cited 3×, 1 in Method
    “Vision-and-language (VL) tasks, such as visual question answering (VQA) antol2015vqa and image-text retrieval lin2014microsoft; plummer2015flickr30k, require an AI system to comprehend both the input image and text conte…”
    From METER · §Introduction
  • Simple end-to-end captioning2022 · cited 3×, 1 in Method
    “For image captioning task, we use three different datasets, including MSCOCO Captions (Lin et al. 2015)44 4 https://cocodataset.org/#home, Flickr30k (Plummer et al. 2016)55 5 http://hockenmaier.cs.illinois.edu/Denotation…”
    From Simple end-to-end captioning · §. Experiment Setup
  • VL-BEiT2022 · cited 2×
    “We conduct vision-language finetuning experiments on the widely used visual question answering [10], natural language for visual reasoning [36] and image-text retrieval [27, 22] tasks.”
    From VL-BEiT · §Experiments
  • BEiT-32022 · cited 2×
    “We evaluate the capabilities of BEiT-3 on the widely used vision-language understanding and generation benchmarks, including visual question answering [19], visual reasoning [49], image-text retrieval [41, 31], and image…”
    From BEiT-3 · §Experiments on Vision and Vision-Language Tasks
Abstract

The Flickr30k dataset has become a standard benchmark for sentence-based image description. This paper presents Flickr30k Entities, which augments the 158k captions from Flickr30k with 244k coreference chains, linking mentions of the same entities across different captions for the same image, and associating them with 276k manually annotated bounding boxes. Such annotations are essential for continued progress in automatic image description and grounded language understanding. They enable us to define a new benchmark for localization of textual entity mentions in an image. We present a strong baseline for this task that combines an image-text embedding, detectors for common objects, a color classifier, and a bias towards selecting larger objects. While our baseline rivals in accuracy more complex state-of-the-art models, we show that its gains cannot be easily parlayed into improvements on such tasks as image-sentence retrieval, thus underlining the limitations of current methods and the need for further research.