Paper Lineage
Esc
training techniqueIntroduced by BLIP-2 · 2023

Two-stage vision-to-language bootstrapping

First align the bridge with vision and text, then teach it to talk to a frozen LLM.

Drafted by AI · not yet reviewed

How this idea evolved

Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.

  1. 2021

    Keep a pretrained LM fixed and teach only a small module to feed it new information.

    Cites · not yet reviewedcited 3× · §Introduction
    “Frozen (Tsimpoukelli et al. 2021), Flamingo (Alayrac et al. 2022)) resort to an image-to-text generation loss, which we show is insufficient to bridge the modality gap.”
    From BLIP-2 · §Introduction
  2. 2023

    First align the bridge with vision and text, then teach it to talk to a frozen LLM.

    Also draws on: Locked image tuning (LiT)

Papers using this

No other paper in this dataset is tagged with it yet.