objectiveIntroduced by VL-BEiT · 2022
Masked vision-language modelling
Masked prediction over image patches and text tokens together, in one shared backbone.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that VL-BEiT cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Shared attention with per-modality feed-forward experts, usable as a dual or fusion encoder.
Feed raw image patches straight into the VL Transformer: no CNN, no detector.
Align unimodal embeddings contrastively before fusing them with cross-attention.
Papers using this
- 2019Unified VLP