Flamingo: a Visual Language Model for Few-Shot Learning
Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Flamingo, a family of Visual Language Models (VLM) with this ability.
Also cited · not yet reviewed (13)
- GPT-32020 · cited 6×, 1 in Method“We evaluate the ability of our models to rapidly adapt to new tasks using in-context learning, analogously to GPT-3 [11], by interleaving support example pairs in the form of (image,text)(image,text) or (video…”From this paper · §Approach
- CLIP2021 · cited 5×, 1 in Method“We pretrain the vision encoder using a contrastive objective on our datasets of image and text pairs, using the two-term contrastive loss from Radford et al. 2021.”From this paper · §Approach
- Chinchilla2022 · cited 5×, 1 in Method“We perform experiments across three models sizes, building on the 1.4B, 7B, and 70B parameter Chinchilla models [42]; calling them respectively Flamingo-3B, Flamingo-9B and Flamingo-80B.”From this paper · §Approach
- VL-T52021 · cited 4×, 1 in Method“We accumulate gradients over all datasets, which we found outperforms a “round-robin” approach [17].”From this paper · §Approach
Show 9 more
- ALIGN2021 · cited 4×, 1 in Method“Another family of vision-language models is based on contrastive learning [2, 85, 50, 146, 82, 74, 5, 140, 57, 138, 49].”From this paper · §Related work
- Simple end-to-end captioning2022 · cited 3×, 1 in Method“In the grafting approach from [68], the frozen LM is used as is with no additional layers inserted, and a stack of interleaved self-attention and cross-attention layers that take the frozen LM output are learnt from scra…”From this paper · §Experiments
- Perceiver2021 · cited 2×, 1 in Method“Similar to Perceiver [48] and DETR [13], we learn a predefined number of latent input queries which are fed to a Transformer and cross-attend to the visual features.”From this paper · §Approach
- VirTex2020 · cited 1×, 1 in Method“In our ablation studies (Section 3.3), we compare the proposed gated xattn-dense layers against recent alternatives [22, 68] and explore the effect of how frequently these additional layers are inserted to trade off betw…”From this paper · §Approach
- Transformer2017 · cited 2דLanguage modelling has recently made substantial progress following the introduction of Transformers [115].”From this paper · §Related work
- BERT2018 · cited 2דThe paradigm of first pretraining on a vast amount of data followed by an adaptation on a downstream task has become standard [75, 32, 52, 44, 23, 87, 108, 11].”From this paper · §Related work
- ViLBERT2019 · cited 2דIn particular, BERT [23] inspired a large body of vision-language work [66, 106, 16, 38, 121, 61, 109, 151, 118, 59, 29, 28, 142, 143, 101, 107].”From this paper · §Related work
- Frozen2021 · cited 2דOne recent line of work [114, 26, 78, 68, 136, 144] proposes to freeze the pretrained LM weights to prevent catastrophic forgetting [71].”From this paper · §Related work
- SimVLM2021 · cited 2דConcurrent works [124, 17, 119, 154, 58] also propose to formulate numerous vision tasks as text generation problems.”From this paper · §Related work
Led to
- Emergent abilities2022 · cited 3דAs a few examples, GPT-3 175B achieved new state of the art on the TriviaQA and PiQA question-answering benchmarks (Brown et al. 2020); PaLM 540B achieved new state of the art on three arithmetic reasoning benchmarks (Ch…”From Emergent abilities · §Discussion
- BLIP-22023 · cited 5דFrozen (Tsimpoukelli et al. 2021), Flamingo (Alayrac et al. 2022)) resort to an image-to-text generation loss, which we show is insufficient to bridge the modality gap.”From BLIP-2 · §Introduction
Abstract
Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Flamingo, a family of Visual Language Models (VLM) with this ability. We propose key architectural innovations to: (i) bridge powerful pretrained vision-only and language-only models, (ii) handle sequences of arbitrarily interleaved visual and textual data, and (iii) seamlessly ingest images or videos as inputs. Thanks to their flexibility, Flamingo models can be trained on large-scale multimodal web corpora containing arbitrarily interleaved text and images, which is key to endow them with in-context few-shot learning capabilities. We perform a thorough evaluation of our models, exploring and measuring their ability to rapidly adapt to a variety of image and video tasks. These include open-ended tasks such as visual question-answering, where the model is prompted with a question which it has to answer; captioning tasks, which evaluate the ability to describe a scene or an event; and close-ended tasks such as multiple-choice visual question-answering. For tasks lying anywhere on this spectrum, a single Flamingo model can achieve a new state of the art with few-shot learning, simply by prompting the model with task-specific examples. On numerous benchmarks, Flamingo outperforms models fine-tuned on thousands of times more task-specific data.