architectureIntroduced by Captions to Visual Concepts · 2014
Visual word detectors for captioning
Detect caption words with multiple-instance learning, then compose sentences with an LM.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that Captions to Visual Concepts cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Score a caption by its TF-IDF n-gram agreement with many human references.
Align image regions with sentence fragments, then generate descriptions with a multimodal RNN.