Multimodal Few-Shot Learning with Frozen Language Models
When trained at sufficient scale, auto-regressive language models exhibit the notable ability to learn a new language task after being prompted with just a few examples. Here, we present a simple, yet effective, approach for transferring this few-shot learning ability to a multimodal setting (vision and language).
Also cited · not yet reviewed (7)
- Prefix-Tuning2021 · cited 4×, 1 in Method“Frozen is a method for grounding a large language model without changing its weights, closely related to prefix tuning [22, 19].”From this paper · §The Frozen Method
- Knowledge in LM parameters2020 · cited 3×, 1 in Method“We use a 7 billion parameter transformer trained on the public dataset C4 [30] – previous work has shown that the multi-billion parameter scale is sufficient to exhibit the key capacities we are interested in studying [2…”From this paper · §The Frozen Method
- Transformer2017 · cited 2×, 1 in Method“Our method starts from a pre-trained deep auto-regressive language model, based on the Transformer architecture [40, 29], which parametrizes a probability distribution over text 𝐲y.”From this paper · §The Frozen Method
- T52019 · cited 2×, 1 in Method“We use a 7 billion parameter transformer trained on the public dataset C4 [30] – previous work has shown that the multi-billion parameter scale is sufficient to exhibit the key capacities we are interested in studying [2…”From this paper · §The Frozen Method
Show 3 more
- Prompt tuning2021 · cited 2×, 1 in Method“Frozen is a method for grounding a large language model without changing its weights, closely related to prefix tuning [22, 19].”From this paper · §The Frozen Method
- SentencePiece2018 · cited 1×, 1 in Method“Text is decomposed into a sequence of discrete tokens 𝐲=y1,y2,…,yLy=y_{1},y_{2},...,y_{L} by the SentencePiece tokenizer [17].”From this paper · §The Frozen Method
- GPT-32020 · cited 4דAuto-regressive transformers have been shown to be very impressive models of natural language [40].”From this paper · §Introduction
Led to
- Flamingo2022 · cited 2דOne recent line of work [114, 26, 78, 68, 136, 144] proposes to freeze the pretrained LM weights to prevent catastrophic forgetting [71].”From Flamingo · §Related work
- BLIP-22023 · cited 3דFrozen (Tsimpoukelli et al. 2021), Flamingo (Alayrac et al. 2022)) resort to an image-to-text generation loss, which we show is insufficient to bridge the modality gap.”From BLIP-2 · §Introduction
Abstract
When trained at sufficient scale, auto-regressive language models exhibit the notable ability to learn a new language task after being prompted with just a few examples. Here, we present a simple, yet effective, approach for transferring this few-shot learning ability to a multimodal setting (vision and language). Using aligned image and caption data, we train a vision encoder to represent each image as a sequence of continuous embeddings, such that a pre-trained, frozen language model prompted with this prefix generates the appropriate caption. The resulting system is a multimodal few-shot learner, with the surprising ability to learn a variety of new tasks when conditioned on examples, represented as a sequence of multiple interleaved image and text embeddings. We demonstrate that it can rapidly learn words for new objects and novel visual categories, do visual question-answering with only a handful of examples, and make use of outside knowledge, by measuring a single model on a variety of established and new benchmarks.