dataIntroduced by Flamingo · 2022
Interleaved image–text sequences
Train and prompt on web pages where images and text alternate, enabling multimodal few-shot learning.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that Flamingo cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Keep a pretrained LM fixed and teach only a small module to feed it new information.
Turn an image into a few embeddings that a frozen LM reads as if they were words.
Papers using this
- 2021Frozen