BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
The cost of vision-and-language pre-training has become increasingly prohibitive due to end-to-end training of large-scale models. This paper proposes BLIP-2, a generic and efficient pre-training strategy that bootstraps vision-language pre-training from off-the-shelf frozen pre-trained image encoders and frozen large language models.
Also cited · not yet reviewed (16)
- BLIP2022 · cited 9×, 3 in Method“Inspired by BLIP (Li et al. 2022), we jointly optimize three pre-training objectives that share the same input format and model parameters.”From this paper · §Method
- ALBEF2021 · cited 7×, 1 in Method“We adopt the hard negative mining strategy from Li et al. 2021; Li et al. 2022 to create informative negative pairs.”From this paper · §Method
- CLIP2021 · cited 5×, 1 in Method“For the frozen image encoder, we explore two state-of-the-art pre-trained vision transformer models: (1) ViT-L/14 from CLIP (Radford et al. 2021) and (2) ViT-g/14 from EVA-CLIP (Fang et al. 2022).”From this paper · §Method
- OPT2022 · cited 3×, 1 in Method“For the frozen language model, we explore the unsupervised-trained OPT model family (Zhang et al. 2022) for decoder-based LLMs, and the instruction-trained FlanT5 model family (Chung et al. 2022) for encoder-decoder-base…”From this paper · §Method
Show 12 more
- Flan-T5 / Flan-PaLM2022 · cited 3×, 1 in Method“For the frozen language model, we explore the unsupervised-trained OPT model family (Zhang et al. 2022) for decoder-based LLMs, and the instruction-trained FlanT5 model family (Chung et al. 2022) for encoder-decoder-base…”From this paper · §Method
- MS COCO2014 · cited 1×, 1 in Method“We use the same pre-training dataset as BLIP with 129M images in total, including COCO (Lin et al. 2014), Visual Genome (Krishna et al. 2017), CC3M (Sharma et al. 2018), CC12M (Changpinyo et al. 2021), SBU (Ordonez et al…”From this paper · §Method
- Visual Genome2016 · cited 1×, 1 in Method“We use the same pre-training dataset as BLIP with 129M images in total, including COCO (Lin et al. 2014), Visual Genome (Krishna et al. 2017), CC3M (Sharma et al. 2018), CC12M (Changpinyo et al. 2021), SBU (Ordonez et al…”From this paper · §Method
- AdamW2017 · cited 1×, 1 in Method“We use the AdamW (Loshchilov & Hutter 2017) optimizer with β1=0.9\beta_{1}=0.9, β1=0.98\beta_{1}=0.98, and a weight decay of 0.05.”From this paper · §Method
- BERT2018 · cited 1×, 1 in Method“We initialize Q-Former with the pre-trained weights of BERTbase (Devlin et al. 2019), whereas the cross-attention layers are randomly initialized.”From this paper · §Method
- UniLM2019 · cited 1×, 1 in Method“We employ a multimodal causal self-attention mask to control query-text interaction, similar to the one used in UniLM (Dong et al. 2019).”From this paper · §Method
- Conceptual 12M2021 · cited 1×, 1 in Method“We use the same pre-training dataset as BLIP with 129M images in total, including COCO (Lin et al. 2014), Visual Genome (Krishna et al. 2017), CC3M (Sharma et al. 2018), CC12M (Changpinyo et al. 2021), SBU (Ordonez et al…”From this paper · §Method
- LAION-400M2021 · cited 1×, 1 in Method“We use the same pre-training dataset as BLIP with 129M images in total, including COCO (Lin et al. 2014), Visual Genome (Krishna et al. 2017), CC3M (Sharma et al. 2018), CC12M (Changpinyo et al. 2021), SBU (Ordonez et al…”From this paper · §Method
- EVA2022 · cited 1×, 1 in Method“For the frozen image encoder, we explore two state-of-the-art pre-trained vision transformer models: (1) ViT-L/14 from CLIP (Radford et al. 2021) and (2) ViT-g/14 from EVA-CLIP (Fang et al. 2022).”From this paper · §Method
- Flamingo2022 · cited 5דFrozen (Tsimpoukelli et al. 2021), Flamingo (Alayrac et al. 2022)) resort to an image-to-text generation loss, which we show is insufficient to bridge the modality gap.”From this paper · §Introduction
- Frozen2021 · cited 3דFrozen (Tsimpoukelli et al. 2021), Flamingo (Alayrac et al. 2022)) resort to an image-to-text generation loss, which we show is insufficient to bridge the modality gap.”From this paper · §Introduction
- BEiT-32022 · cited 3דDepending on the downstream task, different model architectures have been proposed, including the dual-encoder architecture (Radford et al. 2021; Jia et al. 2021), the fusion-encoder architecture (Tan & Bansal 2019; Li e…”From this paper · §Related Work
Led to
Nothing yet. This dataset starts here and was traced backwards. Suggest a paper that builds on BLIP-2 →
Abstract
The cost of vision-and-language pre-training has become increasingly prohibitive due to end-to-end training of large-scale models. This paper proposes BLIP-2, a generic and efficient pre-training strategy that bootstraps vision-language pre-training from off-the-shelf frozen pre-trained image encoders and frozen large language models. BLIP-2 bridges the modality gap with a lightweight Querying Transformer, which is pre-trained in two stages. The first stage bootstraps vision-language representation learning from a frozen image encoder. The second stage bootstraps vision-to-language generative learning from a frozen language model. BLIP-2 achieves state-of-the-art performance on various vision-language tasks, despite having significantly fewer trainable parameters than existing methods. For example, our model outperforms Flamingo80B by 8.7% on zero-shot VQAv2 with 54x fewer trainable parameters. We also demonstrate the model's emerging capabilities of zero-shot image-to-text generation that can follow natural language instructions.