Paper Lineage
Esc
MethodJun 2021arXiv 2106.13884cs.CV

Multimodal Few-Shot Learning with Frozen Language Models

Maria Tsimpoukelli, Jacob Menick, Serkan Cabi and 3 others

When trained at sufficient scale, auto-regressive language models exhibit the notable ability to learn a new language task after being prompted with just a few examples. Here, we present a simple, yet effective, approach for transferring this few-shot learning ability to a multimodal setting (vision and language).

From the abstract

Built on

7 papers · 0 verifiedSee as graph

Also cited · not yet reviewed (7)

  • Prefix-Tuning2021 · cited 4×, 1 in Method
    “Frozen is a method for grounding a large language model without changing its weights, closely related to prefix tuning [22, 19].”
    From this paper · §The Frozen Method
  • Knowledge in LM parameters2020 · cited 3×, 1 in Method
    “We use a 7 billion parameter transformer trained on the public dataset C4 [30] – previous work has shown that the multi-billion parameter scale is sufficient to exhibit the key capacities we are interested in studying [2…”
    From this paper · §The Frozen Method
  • Transformer2017 · cited 2×, 1 in Method
    “Our method starts from a pre-trained deep auto-regressive language model, based on the Transformer architecture [40, 29], which parametrizes a probability distribution over text 𝐲y.”
    From this paper · §The Frozen Method
  • T52019 · cited 2×, 1 in Method
    “We use a 7 billion parameter transformer trained on the public dataset C4 [30] – previous work has shown that the multi-billion parameter scale is sufficient to exhibit the key capacities we are interested in studying [2…”
    From this paper · §The Frozen Method
Show 3 more
  • Prompt tuning2021 · cited 2×, 1 in Method
    “Frozen is a method for grounding a large language model without changing its weights, closely related to prefix tuning [22, 19].”
    From this paper · §The Frozen Method
  • SentencePiece2018 · cited 1×, 1 in Method
    “Text is decomposed into a sequence of discrete tokens 𝐲=y1,y2,…,yLy=y_{1},y_{2},...,y_{L} by the SentencePiece tokenizer [17].”
    From this paper · §The Frozen Method
  • GPT-32020 · cited 4×
    “Auto-regressive transformers have been shown to be very impressive models of natural language [40].”
    From this paper · §Introduction

Led to

  • Flamingo2022 · cited 2×
    “One recent line of work [114, 26, 78, 68, 136, 144] proposes to freeze the pretrained LM weights to prevent catastrophic forgetting [71].”
    From Flamingo · §Related work
  • BLIP-22023 · cited 3×
    “Frozen (Tsimpoukelli et al. 2021), Flamingo (Alayrac et al. 2022)) resort to an image-to-text generation loss, which we show is insufficient to bridge the modality gap.”
    From BLIP-2 · §Introduction
Abstract

When trained at sufficient scale, auto-regressive language models exhibit the notable ability to learn a new language task after being prompted with just a few examples. Here, we present a simple, yet effective, approach for transferring this few-shot learning ability to a multimodal setting (vision and language). Using aligned image and caption data, we train a vision encoder to represent each image as a sequence of continuous embeddings, such that a pre-trained, frozen language model prompted with this prefix generates the appropriate caption. The resulting system is a multimodal few-shot learner, with the surprising ability to learn a variety of new tasks when conditioned on examples, represented as a sequence of multiple interleaved image and text embeddings. We demonstrate that it can rapidly learn words for new objects and novel visual categories, do visual question-answering with only a handful of examples, and make use of outside knowledge, by measuring a single model on a variety of established and new benchmarks.