architectureIntroduced by ViTCAP · 2021
Semantic concept tokens for captioning
Detector-free captioning that predicts semantic concepts from ViT grid features.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that ViTCAP cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Second-order bilinear interactions inside attention for captioning.
Memory-augmented region encoding plus mesh connectivity across encoder layers.
Align image regions with sentence fragments, then generate descriptions with a multimodal RNN.