Neural Baby Talk
We introduce a novel framework for image captioning that can produce natural language explicitly grounded in entities that object detectors find in the image. Our approach reconciles classical slot filling approaches (that are generally better grounded in images) with modern neural captioning approaches (that are generally more natural sounding and accurate).
Also cited · not yet reviewed (9)
- Faster R-CNN2015 · cited 4×, 3 in Method“Given an image 𝑰\bm{I}, and the corresponding caption 𝒚\bm{y}, the candidate grounding regions are obtained by using a pre-trained Faster-RCNN network ren2015faster.”From this paper · §Method
- MS COCO2014 · cited 5×, 2 in Method“COCO lin2014microsoft) do not contain phrase grounding annotations, while some datasets do (e.g.”From this paper · §Method
- Flickr30k Entities2015 · cited 4×, 2 in Method“Another related line of work is on resolving referring expressions kazemzadeh2014referitgame (or description-based object retrieval plummer2015flickr30k; hu2016modeling; hu2016natural; rohrbach2016grounding – given a des…”From this paper · §Related Work
- ResNet2015 · cited 2×, 2 in Method“We use Faster R-CNN ren2015faster with ResNet-101 he2015deep to obtain region proposals for the image.”From this paper · §Method
Show 5 more
- Bottom-Up Top-Down attention2017 · cited 7×, 1 in Method“We use an attention model with two LSTM layers Anderson2017up-down as our base attention model.”From this paper · §Method
- Constrained beam search captioning2016 · cited 5דIn that context, our work is related to prior works on novel object captioning anne2016deep; venugopalan2016captioning; yao2017incorporating; anderson2016guided.”From this paper · §Related Work
- Karpathy visual-semantic alignment2014 · cited 3דFor standard image captioning, we use splits from Karpathy et al. karpathy2015deep on COCO/Flickr30k.”From this paper · §Experimental Results
- Captions to Visual Concepts2014 · cited 2דWhile there are many recent extensions of this basic idea to include attention Xu2015show; fang2015captions; you2016image; yang2016encode; lu2016knowing, it is well-understood that models still lack visual grounding (i.e…”From this paper · §Introduction
- CIDEr2014 · cited 2דWe report results using the COCO captioning evaluation toolkit lin2014microsoft, which reports the widely used automatic evaluation metrics, BLEU papineni2002bleu, METEOR denkowski2014meteor, CIDEr vedantam2015cider and…”From this paper · §Experimental Results
Led to
- nocaps2018 · cited 10×, 3 in Method“To establish the state-of-the-art on our challenging benchmark, we evaluate two of the best performing existing approaches [27, 2] and report their performance based on well-established evaluation metrics – CIDEr [41] an…”From nocaps · §Introduction
- Meshed-Memory Transformer2019 · cited 2דOn the image encoding side, instead, single-layer attention mechanisms have been adopted to incorporate spatial knowledge, initially from a grid of CNN features xu2015show; lu2017knowing; you2016image, and then using ima…”From Meshed-Memory Transformer · §Related work
- Simple end-to-end captioning2022 · cited 2×, 1 in Method“Early works are based on the features extracted by a pre-trained image classification model (Chen et al. 2017; Chen and Zitnick 2014), while the later works (Rennie et al. 2017a; Lu et al. 2017; Anderson et al. 2018; Lu…”From Simple end-to-end captioning · §. Related Work
Abstract
We introduce a novel framework for image captioning that can produce natural language explicitly grounded in entities that object detectors find in the image. Our approach reconciles classical slot filling approaches (that are generally better grounded in images) with modern neural captioning approaches (that are generally more natural sounding and accurate). Our approach first generates a sentence `template' with slot locations explicitly tied to specific image regions. These slots are then filled in by visual concepts identified in the regions by object detectors. The entire architecture (sentence template generation and slot filling with object detectors) is end-to-end differentiable. We verify the effectiveness of our proposed model on different image captioning tasks. On standard image captioning and novel object captioning, our model reaches state-of-the-art on both COCO and Flickr30k datasets. We also demonstrate that our model has unique advantages when the train and test distributions of scene compositions -- and hence language priors of associated captions -- are different. Code has been made available at: https://github.com/jiasenlu/NeuralBabyTalk