Vision-language models on frozen LLMs
Connect a pretrained vision encoder to a frozen large language model through a small trainable bridge.
| Concept | Introduced by | Years | Papers |
|---|---|---|---|
| Frozen language model Keep a pretrained LM fixed and teach only a small module to feed it new information. | Frozen | 2021–2023 | 1 |
| Visual prefix for a frozen LM Turn an image into a few embeddings that a frozen LM reads as if they were words. | Frozen | 2021–2022 | 4 |
| Gated cross-attention layers New cross-attention layers inside a frozen LLM, gated to start as identity. | Flamingo | 2022 | 0 |
| Interleaved image–text sequences Train and prompt on web pages where images and text alternate, enabling multimodal few-shot learning. | Flamingo | 2022 | 1 |
| Perceiver resampler Learned queries cross-attend to many visual features and compress them to a fixed number of tokens. | Flamingo | 2022 | 0 |
| Querying Transformer (Q-Former) A lightweight Transformer whose learned queries extract the most text-relevant features from a frozen image encoder. | BLIP-2 | 2023 | 0 |
| Two-stage vision-to-language bootstrapping First align the bridge with vision and text, then teach it to talk to a frozen LLM. | BLIP-2 | 2023 | 0 |