A Frustratingly Simple Approach for End-to-End Image Captioning
Image Captioning is a fundamental task to join vision and language, concerning about cross-modal understanding and text generation. Recent years witness the emerging attention on image captioning.
Also cited · not yet reviewed (20)
- Oscar2020 · cited 5×, 3 in Method“We follow the suggestion of Li et al. 2020 to evaluate our models with the validation set, containing 4.5k images and 10 captions per image.”From this paper · §. Experiment Setup
- ViTCAP2021 · cited 5×, 3 in Method“Most of the previous works neglect the importance of single-modal generation ability, but train their language decoder from scratch (Xia et al. 2021; Cho et al. 2021; Fang et al. 2021).”From this paper · §. Methodology
- Bottom-Up Top-Down attention2017 · cited 4×, 2 in Method“Early works are based on the features extracted by a pre-trained image classification model (Chen et al. 2017; Chen and Zitnick 2014), while the later works (Rennie et al. 2017a; Lu et al. 2017; Anderson et al. 2018; Lu…”From this paper · §. Related Work
- GCN-LSTM captioning2018 · cited 4×, 2 in Method“Early works are based on the features extracted by a pre-trained image classification model (Chen et al. 2017; Chen and Zitnick 2014), while the later works (Rennie et al. 2017a; Lu et al. 2017; Anderson et al. 2018; Lu…”From this paper · §. Related Work
Show 16 more
- Meshed-Memory Transformer2019 · cited 4×, 2 in Method“Early works are based on the features extracted by a pre-trained image classification model (Chen et al. 2017; Chen and Zitnick 2014), while the later works (Rennie et al. 2017a; Lu et al. 2017; Anderson et al. 2018; Lu…”From this paper · §. Related Work
- X-Linear attention2020 · cited 4×, 2 in Method“Early works are based on the features extracted by a pre-trained image classification model (Chen et al. 2017; Chen and Zitnick 2014), while the later works (Rennie et al. 2017a; Lu et al. 2017; Anderson et al. 2018; Lu…”From this paper · §. Related Work
- CIDEr2014 · cited 2×, 2 in Method“We follow this paradigm to directly optimize the image captioning evaluation metric (CIDEr (Vedantam et al. 2015)) with the widely used self-critical sequence training (SCST) REINFORCE algorithm (Rennie et al. 2017b).”From this paper · §. Methodology
- MS COCO2014 · cited 4×, 1 in Method“For image captioning task, we use three different datasets, including MSCOCO Captions (Lin et al. 2015)44 4 https://cocodataset.org/#home, Flickr30k (Plummer et al. 2016)55 5 http://hockenmaier.cs.illinois.edu/Denotation…”From this paper · §. Experiment Setup
- Flickr30k Entities2015 · cited 3×, 1 in Method“For image captioning task, we use three different datasets, including MSCOCO Captions (Lin et al. 2015)44 4 https://cocodataset.org/#home, Flickr30k (Plummer et al. 2016)55 5 http://hockenmaier.cs.illinois.edu/Denotation…”From this paper · §. Experiment Setup
- ViT2020 · cited 3×, 1 in Method“Visual Encoder.We follow the same setting as the visual transformer (ViT) which recently attracts much attention and shows remarkable performance in the computer vision area (Dosovitskiy et al. 2021; Radford et al. 2021;…”From this paper · §. Methodology
- CLIP2021 · cited 3×, 1 in Method“Visual Encoder.We follow the same setting as the visual transformer (ViT) which recently attracts much attention and shows remarkable performance in the computer vision area (Dosovitskiy et al. 2021; Radford et al. 2021;…”From this paper · §. Methodology
- Neural Baby Talk2018 · cited 2×, 1 in Method“Early works are based on the features extracted by a pre-trained image classification model (Chen et al. 2017; Chen and Zitnick 2014), while the later works (Rennie et al. 2017a; Lu et al. 2017; Anderson et al. 2018; Lu…”From this paper · §. Related Work
- nocaps2018 · cited 2×, 1 in Method“For image captioning task, we use three different datasets, including MSCOCO Captions (Lin et al. 2015)44 4 https://cocodataset.org/#home, Flickr30k (Plummer et al. 2016)55 5 http://hockenmaier.cs.illinois.edu/Denotation…”From this paper · §. Experiment Setup
- BEiT2021 · cited 2×, 1 in Method“Visual Encoder.We follow the same setting as the visual transformer (ViT) which recently attracts much attention and shows remarkable performance in the computer vision area (Dosovitskiy et al. 2021; Radford et al. 2021;…”From this paper · §. Methodology
- MAE2021 · cited 2×, 1 in Method“Visual Encoder.We follow the same setting as the visual transformer (ViT) which recently attracts much attention and shows remarkable performance in the computer vision area (Dosovitskiy et al. 2021; Radford et al. 2021;…”From this paper · §. Methodology
- Karpathy visual-semantic alignment2014 · cited 1×, 1 in Method“We follow the standard Karpathy’s split (Karpathy and Fei-Fei 2015) to split 113.2k/5k/5k and 29.8k/1k/1k images for train/val/test, respectively.”From this paper · §. Experiment Setup
- ResNet2015 · cited 1×, 1 in Method“To achieve our goal, we introduce a frustratingly simple, but effective self-ensemble module to combine the output logits of the cross-modal fusion module and the single-modal language decoder with a residual connection…”From this paper · §. Methodology
- AdamW2017 · cited 1×, 1 in Method“Implementation Details.For all cross-entropy-based experiments, we train our models with the AdamW optimization algorithm (Loshchilov and Hutter 2019), 4k batch size, mixed-precision training and FP16.”From this paper · §. Experiment Setup
- Unified VLP2019 · cited 1×, 1 in Method“(2) Large-scale cross-modal pre-trained models: we include the results of 7 cross-modal pre-trained models, including UVLP (Zhou et al. 2019b), OSCAR (Li et al. 2020), XGPT (Xia et al. 2021), MiniVLM (Wang et al. 2020),…”From this paper · §. Experiment Setup
- METER2021 · cited 3דTo mitigate the need for object detectors, METER (Dou et al. 2021) directly connects RoBERTa (Liu et al. 2019) and CLIP-ViT in a single model.”From this paper · §. Related Work
Led to
- Flamingo2022 · cited 3×, 1 in Method“In the grafting approach from [68], the frozen LM is used as is with no additional layers inserted, and a stack of interleaved self-attention and cross-attention layers that take the frozen LM output are learnt from scra…”From Flamingo · §Experiments
Abstract
Image Captioning is a fundamental task to join vision and language, concerning about cross-modal understanding and text generation. Recent years witness the emerging attention on image captioning. Most of existing works follow a traditional two-stage training paradigm. Before training the captioning models, an extra object detector is utilized to recognize the objects in the image at first. However, they require sizeable datasets with fine-grained object annotation for training the object detector, which is a daunting task. In addition, the errors of the object detectors are easy to propagate to the following captioning models, degenerating models' performance. To alleviate such defects, we propose a frustratingly simple but highly effective end-to-end image captioning framework, Visual Conditioned GPT (VC-GPT), by connecting the pre-trained visual encoder (CLIP-ViT) and language decoder (GPT2). Different from the vanilla connection method that directly inserts the cross-attention modules into GPT2, we come up with a self-ensemble cross-modal fusion mechanism that comprehensively considers both the single- and cross-modal knowledge. As a result, we do not need extra object detectors for model training. Experimental results conducted on three popular image captioning benchmarks (MSCOCO, Flickr30k and NoCaps) demonstrate that our VC-GPT achieves either the best or the second-best performance across all evaluation metrics over extensive baseline systems.