objectiveIntroduced by Karpathy visual-semantic alignment · 2014
Region–word visual-semantic alignment
Align image regions with sentence fragments, then generate descriptions with a multimodal RNN.
Drafted by AI · not yet reviewed
How this idea evolved
No earlier ideas recorded for this concept yet.