Paper Lineage
Esc

End-to-end vision-language pre-training

Vision-language models trained from pixels, without a separate object detector.

Concepts in End-to-end vision-language pre-training
ConceptIntroduced byYearsPapers
Image–text matching (ITM)
A binary classifier decides whether an image and a caption match, trained with hard negatives.
—2019–20215
Grid features instead of regions
Plain convolutional grid features match region features for VQA and run 10× faster.
Grid features for VQA2020–20211
Align before fuse
Align unimodal embeddings contrastively before fusing them with cross-attention.
ALBEF2021–20222
Detector-free patch VL Transformer
Feed raw image patches straight into the VL Transformer: no CNN, no detector.
ViLT2021–20222
Mixture of modality experts
Shared attention with per-modality feed-forward experts, usable as a dual or fusion encoder.
VLMo2021–20222
Momentum distillation
Learn from soft targets produced by a moving-average model, tolerating noisy web pairs.
ALBEF20210
PrefixLM vision-language pre-training
One generative prefix-LM objective on weakly aligned pairs replaces many task-specific losses.
SimVLM2021–20225
Caption bootstrapping (CapFilt)
Generate synthetic captions for web images and filter noisy ones, then retrain on cleaner data.
BLIP20220
Masked vision-language modelling
Masked prediction over image patches and text tokens together, in one shared backbone.
VL-BEiT20221
Multimodal mixture of encoder–decoder
One model that acts as a contrastive encoder, a matching encoder or a caption decoder.
BLIP20220
Multiway Transformer (image as a foreign language)
Treat images as another language and pre-train one multiway model with masked 'language' modelling on all.
BEiT-320220