evaluationIntroduced by Flickr30k Entities · 2015
Phrase grounding
Link each phrase in a caption to the image region it mentions.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that Flickr30k Entities cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
One shared ConvNet does classification, localisation and detection with dense sliding windows.
Papers using this
- 2018Bilinear Attention Networks
- 2018Multi-task hierarchical VL
- 2019ViLBERT
- 2019VisualBERT
- 201912-in-1
- 2019Localized Narratives
- 2022LaMDA