End-to-end vision-language pre-training
Vision-language models trained from pixels, without a separate object detector.
| Concept | Introduced by | Years | Papers |
|---|---|---|---|
| Image–text matching (ITM) A binary classifier decides whether an image and a caption match, trained with hard negatives. | — | 2019–2021 | 5 |
| Grid features instead of regions Plain convolutional grid features match region features for VQA and run 10× faster. | Grid features for VQA | 2020–2021 | 1 |
| Align before fuse Align unimodal embeddings contrastively before fusing them with cross-attention. | ALBEF | 2021–2022 | 2 |
| Detector-free patch VL Transformer Feed raw image patches straight into the VL Transformer: no CNN, no detector. | ViLT | 2021–2022 | 2 |
| Mixture of modality experts Shared attention with per-modality feed-forward experts, usable as a dual or fusion encoder. | VLMo | 2021–2022 | 2 |
| Momentum distillation Learn from soft targets produced by a moving-average model, tolerating noisy web pairs. | ALBEF | 2021 | 0 |
| PrefixLM vision-language pre-training One generative prefix-LM objective on weakly aligned pairs replaces many task-specific losses. | SimVLM | 2021–2022 | 5 |
| Caption bootstrapping (CapFilt) Generate synthetic captions for web images and filter noisy ones, then retrain on cleaner data. | BLIP | 2022 | 0 |
| Masked vision-language modelling Masked prediction over image patches and text tokens together, in one shared backbone. | VL-BEiT | 2022 | 1 |
| Multimodal mixture of encoder–decoder One model that acts as a contrastive encoder, a matching encoder or a caption decoder. | BLIP | 2022 | 0 |
| Multiway Transformer (image as a foreign language) Treat images as another language and pre-train one multiway model with masked 'language' modelling on all. | BEiT-3 | 2022 | 0 |