X-Linear Attention Networks for Image Captioning
Recent progress on fine-grained visual recognition and visual question answering has featured Bilinear Pooling, which effectively models the 2$^{nd}$ order interactions across multi-modal inputs. Nevertheless, there has not been evidence in support of building such interactions concurrently with attention mechanism for image captioning.
Also cited · not yet reviewed (3)
- Bottom-Up Top-Down attention2017 · cited 4דThat prompts the recent state-of-the-art methods anderson2017bottom; Xu:ICML15 to adopt visual attention mechanisms which trigger the interaction between visual content and natural sentence.”From this paper · §Introduction
- Transformer2017 · cited 2דNote that here each key/value is concatenated with the new attended feature, followed with a residual connection and layer normalization as in vaswani2017attention.”From this paper · §X-linear Attention Networks (X-LAN)
- Bilinear Attention Networks2018 · cited 2דRecently, yu2018hierarchical presents a hierarchical bilinear pooling model to aggregate multiple cross-layer bilinear pooling features for fine-grained visual recognition. kim2018bilinear exploits low-rank bilinear pool…”From this paper · §Related Work
Led to
- ViTCAP2021 · cited 5דRecent studies have witnessed its great development which are primarily reflected in the aspects of more advanced cross-modal fusion architectures you2016image; vinyals2015show; xu2015show; rennie2017self; yang2019auto;…”From ViTCAP · §Introduction
- Simple end-to-end captioning2022 · cited 4×, 2 in Method“Early works are based on the features extracted by a pre-trained image classification model (Chen et al. 2017; Chen and Zitnick 2014), while the later works (Rennie et al. 2017a; Lu et al. 2017; Anderson et al. 2018; Lu…”From Simple end-to-end captioning · §. Related Work
Abstract
Recent progress on fine-grained visual recognition and visual question answering has featured Bilinear Pooling, which effectively models the 2$^{nd}$ order interactions across multi-modal inputs. Nevertheless, there has not been evidence in support of building such interactions concurrently with attention mechanism for image captioning. In this paper, we introduce a unified attention block -- X-Linear attention block, that fully employs bilinear pooling to selectively capitalize on visual information or perform multi-modal reasoning. Technically, X-Linear attention block simultaneously exploits both the spatial and channel-wise bilinear attention distributions to capture the 2$^{nd}$ order interactions between the input single-modal or multi-modal features. Higher and even infinity order feature interactions are readily modeled through stacking multiple X-Linear attention blocks and equipping the block with Exponential Linear Unit (ELU) in a parameter-free fashion, respectively. Furthermore, we present X-Linear Attention Networks (dubbed as X-LAN) that novelly integrates X-Linear attention block(s) into image encoder and sentence decoder of image captioning model to leverage higher order intra- and inter-modal interactions. The experiments on COCO benchmark demonstrate that our X-LAN obtains to-date the best published CIDEr performance of 132.0% on COCO Karpathy test split. When further endowing Transformer with X-Linear attention blocks, CIDEr is boosted up to 132.8%. Source code is available at \url{https://github.com/Panda-Peter/image-captioning}.