architectureIntroduced by VinVL · 2021
VL-specific object detector
A bigger detector trained on merged datasets gives much better region features for VL.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that VinVL cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Feed detected object labels as text so they anchor image–text alignment.
Represent an image by features of detected object regions instead of a pixel grid.
One Transformer for both VL understanding and generation, switched by attention masks.