Paper Lineage
Esc
MethodDec 2019arXiv 1912.08226cs.CV

Meshed-Memory Transformer for Image Captioning

Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi, Rita Cucchiara

Transformer-based architectures represent the state of the art in sequence modeling tasks like machine translation and language understanding. Their applicability to multi-modal contexts like image captioning, however, is still largely under-explored.

From the abstract

Built on

8 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (8)

  • “On the image encoding side, instead, single-layer attention mechanisms have been adopted to incorporate spatial knowledge, initially from a grid of CNN features xu2015show; lu2017knowing; you2016image, and then using ima…”
    From this paper · §Related work
  • Transformer2017 · cited 7×
    “The recent advent of fully-attentive models, in which the recurrent relation is abandoned in favour of the use of self-attention, offers unique opportunities in terms of set and sequence modeling performances, as testifi…”
    From this paper · §Introduction
  • GCN-LSTM captioning2018 · cited 3×
    “To further improve the encoding of objects and their relationships, Yao et al. yao2018exploring have proposed to use a graph convolution neural network in the image encoding phase to integrate semantic and spatial relati…”
    From this paper · §Related work
  • CIDEr2014 · cited 2×
    “Following previous works anderson2018bottom, we use the CIDEr-D score as reward, as it well correlates with human judgment vedantam2015cider.”
    From this paper · §Meshed-Memory Transformer
Show 4 more
  • “We follow the splits provided by Karpathy et al. karpathy2015deep, where 5 0005\,000 images are used for validation, 5 0005\,000 for testing and the rest for training.”
    From this paper · §Experiments
  • Neural Baby Talk2018 · cited 2×
    “On the image encoding side, instead, single-layer attention mechanisms have been adopted to incorporate spatial knowledge, initially from a grid of CNN features xu2015show; lu2017knowing; you2016image, and then using ima…”
    From this paper · §Related work
  • BERT2018 · cited 2×
    “The recent advent of fully-attentive models, in which the recurrent relation is abandoned in favour of the use of self-attention, offers unique opportunities in terms of set and sequence modeling performances, as testifi…”
    From this paper · §Introduction
  • nocaps2018 · cited 2×
    “Then, we assess the captioning of novel objects by testing on the recently proposed nocaps dataset agrawal2019nocaps.”
    From this paper · §Experiments

Led to

  • Conceptual 12M2021 · cited 1×, 1 in Method
    “For ic-based pre-training and downstream tasks, we follow the state-of-the-art architecture that heavily rely on self-attention [78] or similar mechanisms [70, 85, 19, 33, 23].”
    From Conceptual 12M · §Evaluating Vision-and-Language Pre-Training Data
  • ViTCAP2021 · cited 5×
    “Recent studies have witnessed its great development which are primarily reflected in the aspects of more advanced cross-modal fusion architectures you2016image; vinyals2015show; xu2015show; rennie2017self; yang2019auto;…”
    From ViTCAP · §Introduction
  • Simple end-to-end captioning2022 · cited 4×, 2 in Method
    “Early works are based on the features extracted by a pre-trained image classification model (Chen et al. 2017; Chen and Zitnick 2014), while the later works (Rennie et al. 2017a; Lu et al. 2017; Anderson et al. 2018; Lu…”
    From Simple end-to-end captioning · §. Related Work
Abstract

Transformer-based architectures represent the state of the art in sequence modeling tasks like machine translation and language understanding. Their applicability to multi-modal contexts like image captioning, however, is still largely under-explored. With the aim of filling this gap, we present M$^2$ - a Meshed Transformer with Memory for Image Captioning. The architecture improves both the image encoding and the language generation steps: it learns a multi-level representation of the relationships between image regions integrating learned a priori knowledge, and uses a mesh-like connectivity at decoding stage to exploit low- and high-level features. Experimentally, we investigate the performance of the M$^2$ Transformer and different fully-attentive models in comparison with recurrent ones. When tested on COCO, our proposal achieves a new state of the art in single-model and ensemble configurations on the "Karpathy" test split and on the online test server. We also assess its performances when describing objects unseen in the training set. Trained models and code for reproducing the experiments are publicly available at: https://github.com/aimagelab/meshed-memory-transformer.