Detector-based vision-language pre-training
BERT-style vision-language models built on object-detector region features (2019–2021).
| Concept | Introduced by | Years | Papers |
|---|---|---|---|
| Object-detector region features Represent an image by features of detected object regions instead of a pixel grid. | Bottom-Up Top-Down attention | 2017–2022 | 9 |
| Masked region modelling Mask image regions and predict their object class or features from the text and other regions. | ViLBERT | 2019–2021 | 6 |
| QA as a pre-training task Add image question answering to the vision-language pre-training mix. | LXMERT | 2019–2021 | 7 |
| Shared encoder–decoder VLP One Transformer for both VL understanding and generation, switched by attention masks. | Unified VLP | 2019 | 0 |
| Single-stream VL Transformer Concatenate words and image regions and run one Transformer over both. | VisualBERT | 2019–2021 | 3 |
| Two-stream co-attention Separate image and text streams exchange information through co-attention layers. | ViLBERT | 2019–2022 | 9 |
| Word–region alignment Explicitly align words to image regions with an optimal-transport objective. | UNITER | 2019–2021 | 7 |
| Object tags as alignment anchors Feed detected object labels as text so they anchor image–text alignment. | Oscar | 2020–2022 | 8 |
| VL-specific object detector A bigger detector trained on merged datasets gives much better region features for VL. | VinVL | 2021–2022 | 4 |