Paper Lineage
Esc

Vision-language models on frozen LLMs

Connect a pretrained vision encoder to a frozen large language model through a small trainable bridge.

Concepts in Vision-language models on frozen LLMs
ConceptIntroduced byYearsPapers
Frozen language model
Keep a pretrained LM fixed and teach only a small module to feed it new information.
Frozen2021–20231
Visual prefix for a frozen LM
Turn an image into a few embeddings that a frozen LM reads as if they were words.
Frozen2021–20224
Gated cross-attention layers
New cross-attention layers inside a frozen LLM, gated to start as identity.
Flamingo20220
Interleaved image–text sequences
Train and prompt on web pages where images and text alternate, enabling multimodal few-shot learning.
Flamingo20221
Perceiver resampler
Learned queries cross-attend to many visual features and compress them to a fixed number of tokens.
Flamingo20220
Querying Transformer (Q-Former)
A lightweight Transformer whose learned queries extract the most text-relevant features from a frozen image encoder.
BLIP-220230
Two-stage vision-to-language bootstrapping
First align the bridge with vision and text, then teach it to talk to a frozen LLM.
BLIP-220230