Paper Lineage
Esc
MethodMay 2018arXiv 1805.07932cs.CV

Bilinear Attention Networks

Jin-Hwa Kim, Jaehyun Jun, Byoung-Tak Zhang

Attention networks in multimodal learning provide an efficient way to utilize given visual information selectively. However, the computational cost to learn attention distributions for every pair of multimodal input channels is prohibitively expensive.

From the abstract

Built on

5 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (5)

  • Flickr30k Entities2015 · cited 6×
    “Moreover, we evaluate the visual grounding of bilinear attention map on Flickr30k Entities [23] outperforming previous methods, along with 25.37% improvement of inference speed taking advantage of the processing of multi…”
    From this paper · §Introduction
  • “We use the image features extracted from bottom-up attention [2].”
    From this paper · §Experiments
  • Visual Genome2016 · cited 3×
    “These features are the output of Faster R-CNN [25], pre-trained using Visual Genome [17].”
    From this paper · §Experiments
  • Faster R-CNN2015 · cited 2×
    “Note that both Query-Adaptive RCNN [10] and our off-the-shelf object detector [2] are based on Faster RCNN [25] and pre-trained on Visual Genome [17].”
    From this paper · §Flickr30k entities results and discussions
Show 1 more
  • VQA v22016 · cited 2×
    “Finally, we validate our proposed method on a large and highly-competitive dataset, VQA 2.0 [8].”
    From this paper · §Introduction

Led to

  • LXMERT2019 · cited 3×, 2 in Method
    “The SotA result is BAN+Counter in Kim et al. 2018, which achieves the best accuracy among other recent works: MFH Yu et al. 2018, Pythia Jiang et al. 2018, DFAF Gao et al. 2019a, and Cycle-Consistency Shah et al. 2019.55…”
    From LXMERT · §Experimental Setup and Results
  • Grid features for VQA2020 · cited 2×
    “In fact, extracting region features is so time-consuming that most state-of-the-art models kim2018bilinear; yu2019deep are directly trained and evaluated on cached visual features.”
    From Grid features for VQA · §Introduction
  • X-Linear attention2020 · cited 2×
    “Recently, yu2018hierarchical presents a hierarchical bilinear pooling model to aggregate multiple cross-layer bilinear pooling features for fine-grained visual recognition. kim2018bilinear exploits low-rank bilinear pool…”
    From X-Linear attention · §Related Work
Abstract

Attention networks in multimodal learning provide an efficient way to utilize given visual information selectively. However, the computational cost to learn attention distributions for every pair of multimodal input channels is prohibitively expensive. To solve this problem, co-attention builds two separate attention distributions for each modality neglecting the interaction between multimodal inputs. In this paper, we propose bilinear attention networks (BAN) that find bilinear attention distributions to utilize given vision-language information seamlessly. BAN considers bilinear interactions among two groups of input channels, while low-rank bilinear pooling extracts the joint representations for each pair of channels. Furthermore, we propose a variant of multimodal residual networks to exploit eight-attention maps of the BAN efficiently. We quantitatively and qualitatively evaluate our model on visual question answering (VQA 2.0) and Flickr30k Entities datasets, showing that BAN significantly outperforms previous methods and achieves new state-of-the-arts on both datasets.