architectureIntroduced by Unified VLP · 2019
Shared encoder–decoder VLP
One Transformer for both VL understanding and generation, switched by attention masks.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that Unified VLP cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Add image question answering to the vision-language pre-training mix.
Mask image regions and predict their object class or features from the text and other regions.
Separate image and text streams exchange information through co-attention layers.
Papers using this
No other paper in this dataset is tagged with it yet.