Paper Lineage
Esc
MethodDec 2021arXiv 2112.05230cs.CV

Injecting Semantic Concepts into End-to-End Image Captioning

Zhiyuan Fang, Jianfeng Wang, Xiaowei Hu and 5 others

Tremendous progress has been made in recent years in developing better image captioning models, yet most of them rely on a separate object detector to extract regional features. Recent vision-language studies are shifting towards the detector-free trend by leveraging grid representations for more flexible model training and faster inference speed.

From the abstract

Built on

17 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (17)

  • MS COCO2014 · cited 5×, 1 in Method
    “In particular, ViTCAP achieves 138.1138.1 CIDEr scores on COCO-caption Karpathy split lin2014microsoft, 108.6108.6 on Google-CC sharma2018conceptual, and 95.495.4 on nocaps agrawal2019nocaps datasets.”
    From this paper · §Introduction
  • Visual Genome2016 · cited 4×, 1 in Method
    “To address the issue, one can simply retrieve the concepts from the open-form captions (e.g., by extracting nouns or adjective words as keywords) as the pseudo ground-truth concepts, or alternatively leverage a pre-train…”
    From this paper · §ViTCAP
  • 12-in-12019 · cited 2×, 1 in Method
    “In total, our pre-training corpus contains 9.99.9M image-text pairs and 4.14.1M independent images, and we follow lu202012 to de-duplicate testing images exist in evaluating datasets.”
    From this paper · §Experiment
  • Conceptual 12M2021 · cited 2×, 1 in Method
    “CC-1212M is the model trained with 1212M image-caption pairs changpinyo2021conceptual.”
    From this paper · §Experiment
Show 13 more
  • Oscar2020 · cited 10×
    “Recent studies have witnessed its great development which are primarily reflected in the aspects of more advanced cross-modal fusion architectures you2016image; vinyals2015show; xu2015show; rennie2017self; yang2019auto;…”
    From this paper · §Introduction
  • “Recent studies have witnessed its great development which are primarily reflected in the aspects of more advanced cross-modal fusion architectures you2016image; vinyals2015show; xu2015show; rennie2017self; yang2019auto;…”
    From this paper · §Introduction
  • ViLT2021 · cited 6×
    “These intermediate operations unavoidably cause training inefficiency and high inference latency at prediction stage kim2021vilt; wang2020minivlm; 2) require box annotations and largely limit the flexibility in training…”
    From this paper · §Introduction
  • Meshed-Memory Transformer2019 · cited 5×
    “Recent studies have witnessed its great development which are primarily reflected in the aspects of more advanced cross-modal fusion architectures you2016image; vinyals2015show; xu2015show; rennie2017self; yang2019auto;…”
    From this paper · §Introduction
  • X-Linear attention2020 · cited 5×
    “Recent studies have witnessed its great development which are primarily reflected in the aspects of more advanced cross-modal fusion architectures you2016image; vinyals2015show; xu2015show; rennie2017self; yang2019auto;…”
    From this paper · §Introduction
  • nocaps2018 · cited 4×
    “In particular, ViTCAP achieves 138.1138.1 CIDEr scores on COCO-caption Karpathy split lin2014microsoft, 108.6108.6 on Google-CC sharma2018conceptual, and 95.495.4 on nocaps agrawal2019nocaps datasets.”
    From this paper · §Introduction
  • ViT2020 · cited 3×
    “ViTCAP is constructed on the basis of a vision transformer dosovitskiy2020image as the stem image encoder.”
    From this paper · §Introduction
  • SimVLM2021 · cited 3×
    “Recent studies have witnessed its great development which are primarily reflected in the aspects of more advanced cross-modal fusion architectures you2016image; vinyals2015show; xu2015show; rennie2017self; yang2019auto;…”
    From this paper · §Introduction
  • ResNet2015 · cited 2×
    “In xu2021e2e, the image is encoded with ResNet he2016deep and the performance (117.3117.3 CIDEr on COCO xu2021e2e) is still far from the state-of-the-art detector-based approach (129.3129.3 CIDEr with VinVL-base zhang202…”
    From this paper · §Introduction
  • BERT2018 · cited 2×
    “The transformer architecture and its instantiations (e.g., BERT devlin2018bert, GPT brown2020language) are well-known for their remarkable performances on natural language processing tasks, which are mostly attributed to…”
    From this paper · §ViTCAP
  • Unified VLP2019 · cited 2×
    “Recent studies have witnessed its great development which are primarily reflected in the aspects of more advanced cross-modal fusion architectures you2016image; vinyals2015show; xu2015show; rennie2017self; yang2019auto;…”
    From this paper · §Introduction
  • Grid features for VQA2020 · cited 2×
    “To address these challenges, there is an emerging trend that more recent works propose to eliminate the detector for the VL pre-training in an end-to-end fashion jiang2020defense; huang2020pixel; kim2021vilt; yan2021grid…”
    From this paper · §Introduction
  • GPT-32020 · cited 2×
    “The transformer architecture and its instantiations (e.g., BERT devlin2018bert, GPT brown2020language) are well-known for their remarkable performances on natural language processing tasks, which are mostly attributed to…”
    From this paper · §ViTCAP

Led to

  • Simple end-to-end captioning2022 · cited 5×, 3 in Method
    “Most of the previous works neglect the importance of single-modal generation ability, but train their language decoder from scratch (Xia et al. 2021; Cho et al. 2021; Fang et al. 2021).”
    From Simple end-to-end captioning · §. Methodology
Abstract

Tremendous progress has been made in recent years in developing better image captioning models, yet most of them rely on a separate object detector to extract regional features. Recent vision-language studies are shifting towards the detector-free trend by leveraging grid representations for more flexible model training and faster inference speed. However, such development is primarily focused on image understanding tasks, and remains less investigated for the caption generation task. In this paper, we are concerned with a better-performing detector-free image captioning model, and propose a pure vision transformer-based image captioning model, dubbed as ViTCAP, in which grid representations are used without extracting the regional features. For improved performance, we introduce a novel Concept Token Network (CTN) to predict the semantic concepts and then incorporate them into the end-to-end captioning. In particular, the CTN is built on the basis of a vision transformer and is designed to predict the concept tokens through a classification task, from which the rich semantic information contained greatly benefits the captioning task. Compared with the previous detector-based models, ViTCAP drastically simplifies the architectures and at the same time achieves competitive performance on various challenging image captioning datasets. In particular, ViTCAP reaches 138.1 CIDEr scores on COCO-caption Karpathy-split, 93.8 and 108.6 CIDEr scores on nocaps, and Google-CC captioning datasets, respectively.