architectureIntroduced by Meshed-Memory Transformer · 2019
Meshed-memory captioning Transformer
Memory-augmented region encoding plus mesh connectivity across encoder layers.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that Meshed-Memory Transformer cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Encode semantic and spatial relations between objects with graph convolutions.
Score a caption by its TF-IDF n-gram agreement with many human references.
Align image regions with sentence fragments, then generate descriptions with a multimodal RNN.