training techniqueIntroduced by BLIP-2 · 2023
Two-stage vision-to-language bootstrapping
First align the bridge with vision and text, then teach it to talk to a frozen LLM.
Drafted by AI · not yet reviewed
How this idea evolved
Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.
- 2021
Keep a pretrained LM fixed and teach only a small module to feed it new information.
Cites · not yet reviewedcited 3× · §Introduction“Frozen (Tsimpoukelli et al. 2021), Flamingo (Alayrac et al. 2022)) resort to an image-to-text generation loss, which we show is insufficient to bridge the modality gap.”
From BLIP-2 · §Introduction - 2023
First align the bridge with vision and text, then teach it to talk to a frozen LLM.
Also draws on: Locked image tuning (LiT)
Papers using this
No other paper in this dataset is tagged with it yet.