BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Vision-Language Pre-training (VLP) has advanced the performance for many vision-language tasks. However, most existing pre-trained models only excel in either understanding-based tasks or generation-based tasks.
Also cited · not yet reviewed (12)
- ALBEF2021 · cited 14×, 5 in Method“Due to the prohibitive expense of acquiring human-annotated texts, most methods (Chen et al. 2020; Li et al. 2020; Li et al. 2021a; Wang et al. 2021; Radford et al. 2021) use image and alt-text pairs crawled from the web…”From this paper · §Related Work
- SimVLM2021 · cited 8×, 1 in Method“Due to the prohibitive expense of acquiring human-annotated texts, most methods (Chen et al. 2020; Li et al. 2020; Li et al. 2021a; Wang et al. 2021; Radford et al. 2021) use image and alt-text pairs crawled from the web…”From this paper · §Related Work
- CLIP2021 · cited 5×, 1 in Method“Due to the prohibitive expense of acquiring human-annotated texts, most methods (Chen et al. 2020; Li et al. 2020; Li et al. 2021a; Wang et al. 2021; Radford et al. 2021) use image and alt-text pairs crawled from the web…”From this paper · §Related Work
- UNITER2019 · cited 3×, 1 in Method“Due to the prohibitive expense of acquiring human-annotated texts, most methods (Chen et al. 2020; Li et al. 2020; Li et al. 2021a; Wang et al. 2021; Radford et al. 2021) use image and alt-text pairs crawled from the web…”From this paper · §Related Work
Show 8 more
- BERT2018 · cited 2×, 1 in Method“The text encoder is the same as BERT (Devlin et al. 2019), where a [CLS] token is appended to the beginning of the text input to summarize the sentence.”From this paper · §Method
- ViT2020 · cited 2×, 1 in Method“We employ a visual transformer (Dosovitskiy et al. 2021) as our image encoder, which divides an input image into patches and encodes them as a sequence of embeddings, with an additional [CLS] token to represent the globa…”From this paper · §Method
- MS COCO2014 · cited 1×, 1 in Method“Due to the prohibitive annotation cost, there exist a limited number of high-quality human-annotated image-text pairs {(Ih,Th)}\{(I_{h},T_{h})\} (e.g., COCO (Lin et al. 2014)).”From this paper · §Method
- ViLT2021 · cited 1×, 1 in Method“Compared to using pre-trained object detectors for visual feature extraction (Chen et al. 2020), using a ViT is more computation-friendly and has been adopted by the more recent methods (Li et al. 2021a; Kim et al. 2021)…”From this paper · §Method
- VL-T52021 · cited 3דThere have been many attempts to unify various vision and language tasks into a single framework (Zhou et al. 2020; Cho et al. 2021; Wang et al. 2021).”From this paper · §Related Work
- Conceptual 12M2021 · cited 3דDue to the prohibitive expense of acquiring human-annotated texts, most methods (Chen et al. 2020; Li et al. 2020; Li et al. 2021a; Wang et al. 2021; Radford et al. 2021) use image and alt-text pairs crawled from the web…”From this paper · §Related Work
- Unified VLP2019 · cited 2דThere have been many attempts to unify various vision and language tasks into a single framework (Zhou et al. 2020; Cho et al. 2021; Wang et al. 2021).”From this paper · §Related Work
- Oscar2020 · cited 2דDue to the prohibitive expense of acquiring human-annotated texts, most methods (Chen et al. 2020; Li et al. 2020; Li et al. 2021a; Wang et al. 2021; Radford et al. 2021) use image and alt-text pairs crawled from the web…”From this paper · §Related Work
Led to
- CoCa2022 · cited 2×, 1 in Method“On the other hand, while many existing methods [32, 33, 35, 30, 36, 37] train model components with multiple stages on various data sources and/or modalities, CoCa is pretrained end-to-end from scratch directly with vari…”From CoCa · §Approach
- BEiT-32022 · cited 1×, 1 in Method“In contrast, previous vision-language models [36, 65, 26, 34, 53, 30, 62] usually employ multiple pretraining tasks, such as image-text contrast, image-text matching, and word-patch/region alignment.”From BEiT-3 · §BEiT-3: A General-Purpose Multimodal Foundation Model
- BLIP-22023 · cited 9×, 3 in Method“Inspired by BLIP (Li et al. 2022), we jointly optimize three pre-training objectives that share the same input format and model parameters.”From BLIP-2 · §Method
Abstract
Vision-Language Pre-training (VLP) has advanced the performance for many vision-language tasks. However, most existing pre-trained models only excel in either understanding-based tasks or generation-based tasks. Furthermore, performance improvement has been largely achieved by scaling up the dataset with noisy image-text pairs collected from the web, which is a suboptimal source of supervision. In this paper, we propose BLIP, a new VLP framework which transfers flexibly to both vision-language understanding and generation tasks. BLIP effectively utilizes the noisy web data by bootstrapping the captions, where a captioner generates synthetic captions and a filter removes the noisy ones. We achieve state-of-the-art results on a wide range of vision-language tasks, such as image-text retrieval (+2.7% in average recall@1), image captioning (+2.8% in CIDEr), and VQA (+1.6% in VQA score). BLIP also demonstrates strong generalization ability when directly transferred to video-language tasks in a zero-shot manner. Code, models, and datasets are released at https://github.com/salesforce/BLIP.