Paper Lineage
Esc
MethodDec 2016arXiv 1612.00576cs.CV

Guided Open Vocabulary Image Captioning with Constrained Beam Search

Peter Anderson, Basura Fernando, Mark Johnson, Stephen Gould

Existing image captioning models do not generalize well to out-of-domain images containing novel scenes or objects. This limitation severely hinders the use of these models in real world applications dealing with images in the wild.

From the abstract

Built on

6 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (6)

  • ResNet2015 · cited 3×, 1 in Method
    “For this task we use the ResNet-50 He et al. 2016 CNN, and train the base model on a combined training set containing 155k images comprised of the MSCOCO Chen et al. 2015 training and validation datasets, and the full Fl…”
    From this paper · §Experiments
  • Karpathy visual-semantic alignment2014 · cited 2×, 1 in Method
    “Our approach to out-of-domain image captioning could be applied to any existing CNN-RNN captioning model that can be decoding using beam search, e.g., Donahue et al. 2015; Mao et al. 2015; Karpathy and Fei-Fei 2015; Viny…”
    From this paper · §Approach
  • VGG2014 · cited 1×, 1 in Method
    “For the CNN component of the model, we evaluate using the 16-layer VGG Simonyan and Zisserman 2015 model and the 50-layer Residual Net He et al. 2016, pretrained on ILSVRC-2012 Russakovsky et al. 2015 in both cases.”
    From this paper · §Approach
  • MS COCO2014 · cited 2×
    “Recently, models incorporating recurrent neural networks (RNNs) have demonstrated promising results on this challenging task Vinyals et al. 2015; Fang et al. 2015; Devlin et al. 2015, leveraging new benchmark datasets su…”
    From this paper · §Introduction
Show 2 more
  • “More generally, the effectiveness of incorporating semantic attributes (i.e., image tags) into caption model training for in-domain data has been established by several works Fang et al. 2015; Wu et al. 2016; Elliot and…”
    From this paper · §Related Work
  • COCO Captions2015 · cited 2×
    “For this task we use the ResNet-50 He et al. 2016 CNN, and train the base model on a combined training set containing 155k images comprised of the MSCOCO Chen et al. 2015 training and validation datasets, and the full Fl…”
    From this paper · §Experiments

Led to

  • Neural Baby Talk2018 · cited 5×
    “In that context, our work is related to prior works on novel object captioning anne2016deep; venugopalan2016captioning; yao2017incorporating; anderson2016guided.”
    From Neural Baby Talk · §Related Work
  • nocaps2018 · cited 10×, 2 in Method
    “To establish the state-of-the-art on our challenging benchmark, we evaluate two of the best performing existing approaches [27, 2] and report their performance based on well-established evaluation metrics – CIDEr [41] an…”
    From nocaps · §Introduction
  • Conceptual 12M2021 · cited 2×
    “Addressing long-tail distributions of visual concepts is an important component of V+L systems that generalize, as long and free-form texts exhibit a large number of compositional, fine-grained categories [89, 54, 19].”
    From Conceptual 12M · §Related Work
Abstract

Existing image captioning models do not generalize well to out-of-domain images containing novel scenes or objects. This limitation severely hinders the use of these models in real world applications dealing with images in the wild. We address this problem using a flexible approach that enables existing deep captioning architectures to take advantage of image taggers at test time, without re-training. Our method uses constrained beam search to force the inclusion of selected tag words in the output, and fixed, pretrained word embeddings to facilitate vocabulary expansion to previously unseen tag words. Using this approach we achieve state of the art results for out-of-domain captioning on MSCOCO (and improved results for in-domain captioning). Perhaps surprisingly, our results significantly outperform approaches that incorporate the same tag predictions into the learning algorithm. We also show that we can significantly improve the quality of generated ImageNet captions by leveraging ground-truth labels.