Paper Lineage
Esc
MethodJan 2022arXiv 2201.12723cs.CV

A Frustratingly Simple Approach for End-to-End Image Captioning

Ziyang Luo, Yadong Xi, Rongsheng Zhang, Jing Ma

Image Captioning is a fundamental task to join vision and language, concerning about cross-modal understanding and text generation. Recent years witness the emerging attention on image captioning.

From the abstract

Built on

20 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (20)

  • Oscar2020 · cited 5×, 3 in Method
    “We follow the suggestion of Li et al. 2020 to evaluate our models with the validation set, containing 4.5k images and 10 captions per image.”
    From this paper · §. Experiment Setup
  • ViTCAP2021 · cited 5×, 3 in Method
    “Most of the previous works neglect the importance of single-modal generation ability, but train their language decoder from scratch (Xia et al. 2021; Cho et al. 2021; Fang et al. 2021).”
    From this paper · §. Methodology
  • Bottom-Up Top-Down attention2017 · cited 4×, 2 in Method
    “Early works are based on the features extracted by a pre-trained image classification model (Chen et al. 2017; Chen and Zitnick 2014), while the later works (Rennie et al. 2017a; Lu et al. 2017; Anderson et al. 2018; Lu…”
    From this paper · §. Related Work
  • GCN-LSTM captioning2018 · cited 4×, 2 in Method
    “Early works are based on the features extracted by a pre-trained image classification model (Chen et al. 2017; Chen and Zitnick 2014), while the later works (Rennie et al. 2017a; Lu et al. 2017; Anderson et al. 2018; Lu…”
    From this paper · §. Related Work
Show 16 more
  • Meshed-Memory Transformer2019 · cited 4×, 2 in Method
    “Early works are based on the features extracted by a pre-trained image classification model (Chen et al. 2017; Chen and Zitnick 2014), while the later works (Rennie et al. 2017a; Lu et al. 2017; Anderson et al. 2018; Lu…”
    From this paper · §. Related Work
  • X-Linear attention2020 · cited 4×, 2 in Method
    “Early works are based on the features extracted by a pre-trained image classification model (Chen et al. 2017; Chen and Zitnick 2014), while the later works (Rennie et al. 2017a; Lu et al. 2017; Anderson et al. 2018; Lu…”
    From this paper · §. Related Work
  • CIDEr2014 · cited 2×, 2 in Method
    “We follow this paradigm to directly optimize the image captioning evaluation metric (CIDEr (Vedantam et al. 2015)) with the widely used self-critical sequence training (SCST) REINFORCE algorithm (Rennie et al. 2017b).”
    From this paper · §. Methodology
  • MS COCO2014 · cited 4×, 1 in Method
    “For image captioning task, we use three different datasets, including MSCOCO Captions (Lin et al. 2015)44 4 https://cocodataset.org/#home, Flickr30k (Plummer et al. 2016)55 5 http://hockenmaier.cs.illinois.edu/Denotation…”
    From this paper · §. Experiment Setup
  • Flickr30k Entities2015 · cited 3×, 1 in Method
    “For image captioning task, we use three different datasets, including MSCOCO Captions (Lin et al. 2015)44 4 https://cocodataset.org/#home, Flickr30k (Plummer et al. 2016)55 5 http://hockenmaier.cs.illinois.edu/Denotation…”
    From this paper · §. Experiment Setup
  • ViT2020 · cited 3×, 1 in Method
    “Visual Encoder.We follow the same setting as the visual transformer (ViT) which recently attracts much attention and shows remarkable performance in the computer vision area (Dosovitskiy et al. 2021; Radford et al. 2021;…”
    From this paper · §. Methodology
  • CLIP2021 · cited 3×, 1 in Method
    “Visual Encoder.We follow the same setting as the visual transformer (ViT) which recently attracts much attention and shows remarkable performance in the computer vision area (Dosovitskiy et al. 2021; Radford et al. 2021;…”
    From this paper · §. Methodology
  • Neural Baby Talk2018 · cited 2×, 1 in Method
    “Early works are based on the features extracted by a pre-trained image classification model (Chen et al. 2017; Chen and Zitnick 2014), while the later works (Rennie et al. 2017a; Lu et al. 2017; Anderson et al. 2018; Lu…”
    From this paper · §. Related Work
  • nocaps2018 · cited 2×, 1 in Method
    “For image captioning task, we use three different datasets, including MSCOCO Captions (Lin et al. 2015)44 4 https://cocodataset.org/#home, Flickr30k (Plummer et al. 2016)55 5 http://hockenmaier.cs.illinois.edu/Denotation…”
    From this paper · §. Experiment Setup
  • BEiT2021 · cited 2×, 1 in Method
    “Visual Encoder.We follow the same setting as the visual transformer (ViT) which recently attracts much attention and shows remarkable performance in the computer vision area (Dosovitskiy et al. 2021; Radford et al. 2021;…”
    From this paper · §. Methodology
  • MAE2021 · cited 2×, 1 in Method
    “Visual Encoder.We follow the same setting as the visual transformer (ViT) which recently attracts much attention and shows remarkable performance in the computer vision area (Dosovitskiy et al. 2021; Radford et al. 2021;…”
    From this paper · §. Methodology
  • Karpathy visual-semantic alignment2014 · cited 1×, 1 in Method
    “We follow the standard Karpathy’s split (Karpathy and Fei-Fei 2015) to split 113.2k/5k/5k and 29.8k/1k/1k images for train/val/test, respectively.”
    From this paper · §. Experiment Setup
  • ResNet2015 · cited 1×, 1 in Method
    “To achieve our goal, we introduce a frustratingly simple, but effective self-ensemble module to combine the output logits of the cross-modal fusion module and the single-modal language decoder with a residual connection…”
    From this paper · §. Methodology
  • AdamW2017 · cited 1×, 1 in Method
    “Implementation Details.For all cross-entropy-based experiments, we train our models with the AdamW optimization algorithm (Loshchilov and Hutter 2019), 4k batch size, mixed-precision training and FP16.”
    From this paper · §. Experiment Setup
  • Unified VLP2019 · cited 1×, 1 in Method
    “(2) Large-scale cross-modal pre-trained models: we include the results of 7 cross-modal pre-trained models, including UVLP (Zhou et al. 2019b), OSCAR (Li et al. 2020), XGPT (Xia et al. 2021), MiniVLM (Wang et al. 2020),…”
    From this paper · §. Experiment Setup
  • METER2021 · cited 3×
    “To mitigate the need for object detectors, METER (Dou et al. 2021) directly connects RoBERTa (Liu et al. 2019) and CLIP-ViT in a single model.”
    From this paper · §. Related Work

Led to

  • Flamingo2022 · cited 3×, 1 in Method
    “In the grafting approach from [68], the frozen LM is used as is with no additional layers inserted, and a stack of interleaved self-attention and cross-attention layers that take the frozen LM output are learnt from scra…”
    From Flamingo · §Experiments
Abstract

Image Captioning is a fundamental task to join vision and language, concerning about cross-modal understanding and text generation. Recent years witness the emerging attention on image captioning. Most of existing works follow a traditional two-stage training paradigm. Before training the captioning models, an extra object detector is utilized to recognize the objects in the image at first. However, they require sizeable datasets with fine-grained object annotation for training the object detector, which is a daunting task. In addition, the errors of the object detectors are easy to propagate to the following captioning models, degenerating models' performance. To alleviate such defects, we propose a frustratingly simple but highly effective end-to-end image captioning framework, Visual Conditioned GPT (VC-GPT), by connecting the pre-trained visual encoder (CLIP-ViT) and language decoder (GPT2). Different from the vanilla connection method that directly inserts the cross-attention modules into GPT2, we come up with a self-ensemble cross-modal fusion mechanism that comprehensively considers both the single- and cross-modal knowledge. As a result, we do not need extra object detectors for model training. Experimental results conducted on three popular image captioning benchmarks (MSCOCO, Flickr30k and NoCaps) demonstrate that our VC-GPT achieves either the best or the second-best performance across all evaluation metrics over extensive baseline systems.