Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
Top-down visual attention mechanisms have been used extensively in image captioning and visual question answering (VQA) to enable deeper image understanding through fine-grained analysis and even multiple steps of reasoning. In this work, we propose a combined bottom-up and top-down attention mechanism that enables attention to be calculated at the level of objects and other salient image regions.
Also cited · not yet reviewed (6)
- Visual Genome2016 · cited 3×, 2 in Method“We then train on Visual Genome krishnavisualgenome data.”From this paper · §Approach
- Faster R-CNN2015 · cited 3×, 1 in Method“Practically, we implement bottom-up attention using Faster R-CNN faster_rcnn, which represents a natural expression of a bottom-up attention mechanism.”From this paper · §Introduction
- CIDEr2014 · cited 2×, 1 in Method“To evaluate caption quality, we use the standard automatic evaluation metrics, namely SPICE spice2016, CIDEr Vedantam2015, METEOR meteor-wmt:2014, ROUGE-L Lin2004 and BLEU Papineni2002.”From this paper · §Evaluation
- ResNet2015 · cited 2×, 1 in Method“In this work, we use Faster R-CNN in conjunction with the ResNet-101 he2015deep CNN.”From this paper · §Approach
Show 2 more
- MS COCO2014 · cited 2דAs approximately 51K Visual Genome images are also found in the MSCOCO captions dataset Lin2014, we are careful to avoid contamination of our MSCOCO validation and test sets.”From this paper · §Evaluation
- VQA v22016 · cited 2דProblems combining image and language understanding such as image captioning Chen2015 and visual question answering (VQA) balanced_vqa_v2 continue to inspire considerable research at the boundary of computer vision and n…”From this paper · §Introduction
Led to
- Neural Baby Talk2018 · cited 7×, 1 in Method“We use an attention model with two LSTM layers Anderson2017up-down as our base attention model.”From Neural Baby Talk · §Method
- Bilinear Attention Networks2018 · cited 5דWe use the image features extracted from bottom-up attention [2].”From Bilinear Attention Networks · §Experiments
- nocaps2018 · cited 5×, 1 in Method“To provide an initial measure of the state-of-the-art on nocaps, we extend and present results for two contemporary approaches to novel object captioning – Neural Baby Talk (NBT) [27] and Constrained Beam Search (CBS) [2…”From nocaps · §Experiments
- ViLBERT2019 · cited 3×, 1 in Method“We use Faster R-CNN [31] (with ResNet-101 [11] backbone) pretrained on the Visual Genome dataset [16] (see [30] for details) to extract region features.”From ViLBERT · §Experimental Settings
- VisualBERT2019 · cited 2דVarious tasks such as visual question answering (Antol et al. 2015; Goyal et al. 2017), textual grounding (Kazemzadeh et al. 2014; Plummer et al. 2015), and visual reasoning (Suhr et al. 2019; Zellers et al. 2019) have b…”From VisualBERT · §Related Work
- LXMERT2019 · cited 7×, 4 in Method“Instead of using the feature map output by a convolutional neural network, we follow Anderson et al. 2018 in taking the features of detected objects as the embeddings of images.”From LXMERT · §Model Architecture
- VL-BERT2019 · cited 3דVisual content embedding is produced by Faster R-CNN + ResNet-101, initialized from parameters pre-trained on Visual Genome (Krishna et al. 2017) for object detection (see BUTD (Anderson et al. 2018)).”From VL-BERT · §Experiment
- Decoupled box proposals captioning2019 · cited 6×, 2 in Method“We reimplement the Faster R-CNN model, training it to predict both 1,600 object and 400 attribute labels in Visual Genome Krishna et al. 2017, following the standard setting from Anderson et al. 2018.”From Decoupled box proposals captioning · §Features and Experimental Setup
- 12-in-12019 · cited 2×, 2 in Method“At the interface level, ViLBERT takes as input an image II and text segment QQ represented as the sequence {\{IMG,v1,…,v𝒯,,v_{1},\dots,v_{T}, CLS, w1,…,wT,w_{1},\dots,w_{T}, SEP}\} where {vi}i=1𝒯\{v_{i}\}_{i=1}^{T} are…”From 12-in-1 · §Approach
- Meshed-Memory Transformer2019 · cited 9דOn the image encoding side, instead, single-layer attention mechanisms have been adopted to incorporate spatial knowledge, initially from a grid of CNN features xu2015show; lu2017knowing; you2016image, and then using ima…”From Meshed-Memory Transformer · §Related work
- Grid features for VQA2020 · cited 17דAfter the introduction of deep learning donahue2015long; vinyals2015show and attention mechanisms xu2015show; yang2016stacked to multi-modal vision and language research, perhaps one of the most significant developments…”From Grid features for VQA · §Introduction
- X-Linear attention2020 · cited 4דThat prompts the recent state-of-the-art methods anderson2017bottom; Xu:ICML15 to adopt visual attention mechanisms which trigger the interaction between visual content and natural sentence.”From X-Linear attention · §Introduction
- VILLA2020 · cited 3דFor text modality, we add adversarial perturbations to word embeddings [45, 86, 24].”From VILLA · §Introduction
- VL-BERT meta-analysis2020 · cited 1×, 1 in Method“Our models are trained with 3636 regions of interest extracted by a Faster R-CNN with a ResNet-101101 backbone (Anderson et al. 2018).”From VL-BERT meta-analysis · §Experimental Setup
- VinVL2021 · cited 18דThe output of the 𝐕𝐢𝐬𝐢𝐨𝐧Vision module consists of 𝒒{\boldsymbol{q}} and 𝒗\boldsymbol{v}. 𝒒{\boldsymbol{q}} is the semantic representation of the image, such as tags or detected objects, and 𝒗\boldsymbol{v} the…”From VinVL · §Improving Vision (V) in Vision Language (VL)
- VL-T52021 · cited 3×, 1 in Method“Following this success, image+text pretraining models (Lu et al. 2019; Tan & Bansal 2019; Chen et al. 2020; Huang et al. 2020; Li et al. 2020b; Cho et al. 2020; Radford et al. 2021; Zhang et al. 2021) and video+text pret…”From VL-T5 · §Related Works
- ViLT2021 · cited 3×, 2 in Method“Most VLP models employ an object detector pre-trained on the Visual Genome dataset (Krishna et al. 2017) annotated with 1,600 object classes and 400 attribute classes as in Anderson et al. 2018.”From ViLT · §Introduction
- Conceptual 12M2021 · cited 1×, 1 in Method“We use Graph-RISE [37, 38] to featurize the entire image.”From Conceptual 12M · §Evaluating Vision-and-Language Pre-Training Data
- Florence2021 · cited 1×, 1 in Method“Thus, the object detector has been a de facto tool for image feature extraction, followed by a fusion network for prediction in many works (Anderson et al. 2018; Li et al. 2020; Zhang et al. 2021b; Wang et al. 2020; Fang…”From Florence · §Approach
- ViTCAP2021 · cited 7דRecent studies have witnessed its great development which are primarily reflected in the aspects of more advanced cross-modal fusion architectures you2016image; vinyals2015show; xu2015show; rennie2017self; yang2019auto;…”From ViTCAP · §Introduction
- Simple end-to-end captioning2022 · cited 4×, 2 in Method“Early works are based on the features extracted by a pre-trained image classification model (Chen et al. 2017; Chen and Zitnick 2014), while the later works (Rennie et al. 2017a; Lu et al. 2017; Anderson et al. 2018; Lu…”From Simple end-to-end captioning · §. Related Work
Abstract
Top-down visual attention mechanisms have been used extensively in image captioning and visual question answering (VQA) to enable deeper image understanding through fine-grained analysis and even multiple steps of reasoning. In this work, we propose a combined bottom-up and top-down attention mechanism that enables attention to be calculated at the level of objects and other salient image regions. This is the natural basis for attention to be considered. Within our approach, the bottom-up mechanism (based on Faster R-CNN) proposes image regions, each with an associated feature vector, while the top-down mechanism determines feature weightings. Applying this approach to image captioning, our results on the MSCOCO test server establish a new state-of-the-art for the task, achieving CIDEr / SPICE / BLEU-4 scores of 117.9, 21.5 and 36.9, respectively. Demonstrating the broad applicability of the method, applying the same approach to VQA we obtain first place in the 2017 VQA Challenge.