Paper Lineage
Esc

Detector-based vision-language pre-training

BERT-style vision-language models built on object-detector region features (2019–2021).

Concepts in Detector-based vision-language pre-training
ConceptIntroduced byYearsPapers
Object-detector region features
Represent an image by features of detected object regions instead of a pixel grid.
Bottom-Up Top-Down attention2017–20229
Masked region modelling
Mask image regions and predict their object class or features from the text and other regions.
ViLBERT2019–20216
QA as a pre-training task
Add image question answering to the vision-language pre-training mix.
LXMERT2019–20217
Shared encoder–decoder VLP
One Transformer for both VL understanding and generation, switched by attention masks.
Unified VLP20190
Single-stream VL Transformer
Concatenate words and image regions and run one Transformer over both.
VisualBERT2019–20213
Two-stream co-attention
Separate image and text streams exchange information through co-attention layers.
ViLBERT2019–20229
Word–region alignment
Explicitly align words to image regions with an optimal-transport objective.
UNITER2019–20217
Object tags as alignment anchors
Feed detected object labels as text so they anchor image–text alignment.
Oscar2020–20228
VL-specific object detector
A bigger detector trained on merged datasets gives much better region features for VL.
VinVL2021–20224