Paper Lineage
Esc
MethodJan 2020arXiv 2001.03615cs.CV

In Defense of Grid Features for Visual Question Answering

Huaizu Jiang, Ishan Misra, Marcus Rohrbach and 2 others

Popularized as 'bottom-up' attention, bounding box (or region) based visual features have recently surpassed vanilla grid-based convolutional features as the de facto standard for vision and language tasks like visual question answering (VQA). g.

From the abstract

Built on

12 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (12)

  • “After the introduction of deep learning donahue2015long; vinyals2015show and attention mechanisms xu2015show; yang2016stacked to multi-modal vision and language research, perhaps one of the most significant developments…”
    From this paper · §Introduction
  • Faster R-CNN2015 · cited 6×
    “Unlike normal attention that uses ‘top-down’ linguistic inputs to focus on specific parts of the visual input, bottom-up attention uses pre-trained object detectors ren2015faster to identify salient regions based solely…”
    From this paper · §Introduction
  • ResNet2015 · cited 6×
    “As a result, images are represented by a collection of bounding box or region11 1 We use the terms ‘region’ and ‘bounding box’ interchangeably.-based features anderson2018bottom; teney2018tips–in contrast to vanilla grid…”
    From this paper · §Introduction
  • Visual Genome2016 · cited 4×
    “In fact, our ablative analysis suggests that the key factors which contributed to the high accuracy of existing bottom-up attention features are: 1) the large-scale object and attribute annotations collected in the Visua…”
    From this paper · §Introduction
Show 8 more
  • VQA v22016 · cited 4×
    “It is also worth noting that while region features are effective on benchmarks like VQA antol2015vqa; goyal2017making and COCO captions chen2015microsoft, for benchmarks that diagnose a model’s reasoning abilities when a…”
    From this paper · §Related Work
  • MS COCO2014 · cited 3×
    “Second, while our 1×\times1 RoIPool-based variant hurts the object detection performance (average precision lin2014microsoft on VG drops from 4.07 to 2.90), it helps VQA – boosting the accuracy by 0.73% (row 3 & 4) and a…”
    From this paper · §Main Comparison: Regions vs. Grids
  • COCO Captions2015 · cited 3×
    “Through a comprehensive set of experiments, we verified that our observations generalize across different network backbones, different VQA models jiang2018pythia; yu2019deep, different VQA benchmarks antol2015vqa; gurari…”
    From this paper · §Introduction
  • VGG2014 · cited 2×
    “As a result, images are represented by a collection of bounding box or region11 1 We use the terms ‘region’ and ‘bounding box’ interchangeably.-based features anderson2018bottom; teney2018tips–in contrast to vanilla grid…”
    From this paper · §Introduction
  • YFCC100M2015 · cited 2×
    “For classification, we include a model trained on YFCC thomee2016yfcc100m, which has 92M images with image tags.”
    From this paper · §Why do Our Grid Features Work?
  • VQA2015 · cited 2×
    “Through a comprehensive set of experiments, we verified that our observations generalize across different network backbones, different VQA models jiang2018pythia; yu2019deep, different VQA benchmarks antol2015vqa; gurari…”
    From this paper · §Introduction
  • “In fact, extracting region features is so time-consuming that most state-of-the-art models kim2018bilinear; yu2019deep are directly trained and evaluated on cached visual features.”
    From this paper · §Introduction
  • ViLBERT2019 · cited 2×
    “As these separately trained features may not be optimal for joint vision and language understanding, a recent hot topic is to develop jointly pre-trained models li2019visualbert; lu2019vilbert; tan2019lxmert; su2019vl; z…”
    From this paper · §Related Work

Led to

  • VinVL2021 · cited 3×
    “Although [23] shows that the FPN model outperforms the C4 model for object detection, recent studies [14] demonstrate that FPN does not provide more effective region features for VL tasks than C4, which is also confirmed…”
    From VinVL · §Improving Vision (V) in Vision Language (VL)
  • ViLT2021 · cited 2×, 2 in Method
    “Applying NMS to each and every class becomes a major runtime bottleneck with a large number of classes, e.g. 1.6K in the VG dataset (Jiang et al. 2020).”
    From ViLT · §Background
  • ViTCAP2021 · cited 2×
    “To address these challenges, there is an emerging trend that more recent works propose to eliminate the detector for the VL pre-training in an end-to-end fashion jiang2020defense; huang2020pixel; kim2021vilt; yan2021grid…”
    From ViTCAP · §Introduction
Abstract

Popularized as 'bottom-up' attention, bounding box (or region) based visual features have recently surpassed vanilla grid-based convolutional features as the de facto standard for vision and language tasks like visual question answering (VQA). However, it is not clear whether the advantages of regions (e.g. better localization) are the key reasons for the success of bottom-up attention. In this paper, we revisit grid features for VQA, and find they can work surprisingly well - running more than an order of magnitude faster with the same accuracy (e.g. if pre-trained in a similar fashion). Through extensive experiments, we verify that this observation holds true across different VQA models (reporting a state-of-the-art accuracy on VQA 2.0 test-std, 72.71), datasets, and generalizes well to other tasks like image captioning. As grid features make the model design and training process much simpler, this enables us to train them end-to-end and also use a more flexible network design. We learn VQA models end-to-end, from pixels directly to answers, and show that strong performance is achievable without using any region annotations in pre-training. We hope our findings help further improve the scientific understanding and the practical application of VQA. Code and features will be made available.