Bilinear Attention Networks
Attention networks in multimodal learning provide an efficient way to utilize given visual information selectively. However, the computational cost to learn attention distributions for every pair of multimodal input channels is prohibitively expensive.
Also cited · not yet reviewed (5)
- Flickr30k Entities2015 · cited 6דMoreover, we evaluate the visual grounding of bilinear attention map on Flickr30k Entities [23] outperforming previous methods, along with 25.37% improvement of inference speed taking advantage of the processing of multi…”From this paper · §Introduction
- Bottom-Up Top-Down attention2017 · cited 5דWe use the image features extracted from bottom-up attention [2].”From this paper · §Experiments
- Visual Genome2016 · cited 3דThese features are the output of Faster R-CNN [25], pre-trained using Visual Genome [17].”From this paper · §Experiments
- Faster R-CNN2015 · cited 2דNote that both Query-Adaptive RCNN [10] and our off-the-shelf object detector [2] are based on Faster RCNN [25] and pre-trained on Visual Genome [17].”From this paper · §Flickr30k entities results and discussions
Show 1 more
- VQA v22016 · cited 2דFinally, we validate our proposed method on a large and highly-competitive dataset, VQA 2.0 [8].”From this paper · §Introduction
Led to
- LXMERT2019 · cited 3×, 2 in Method“The SotA result is BAN+Counter in Kim et al. 2018, which achieves the best accuracy among other recent works: MFH Yu et al. 2018, Pythia Jiang et al. 2018, DFAF Gao et al. 2019a, and Cycle-Consistency Shah et al. 2019.55…”From LXMERT · §Experimental Setup and Results
- Grid features for VQA2020 · cited 2דIn fact, extracting region features is so time-consuming that most state-of-the-art models kim2018bilinear; yu2019deep are directly trained and evaluated on cached visual features.”From Grid features for VQA · §Introduction
- X-Linear attention2020 · cited 2דRecently, yu2018hierarchical presents a hierarchical bilinear pooling model to aggregate multiple cross-layer bilinear pooling features for fine-grained visual recognition. kim2018bilinear exploits low-rank bilinear pool…”From X-Linear attention · §Related Work
Abstract
Attention networks in multimodal learning provide an efficient way to utilize given visual information selectively. However, the computational cost to learn attention distributions for every pair of multimodal input channels is prohibitively expensive. To solve this problem, co-attention builds two separate attention distributions for each modality neglecting the interaction between multimodal inputs. In this paper, we propose bilinear attention networks (BAN) that find bilinear attention distributions to utilize given vision-language information seamlessly. BAN considers bilinear interactions among two groups of input channels, while low-rank bilinear pooling extracts the joint representations for each pair of channels. Furthermore, we propose a variant of multimodal residual networks to exploit eight-attention maps of the BAN efficiently. We quantitatively and qualitatively evaluate our model on visual question answering (VQA 2.0) and Flickr30k Entities datasets, showing that BAN significantly outperforms previous methods and achieves new state-of-the-arts on both datasets.