Paper Lineage
Esc
architectureIntroduced by BLIP-2 · 2023

Querying Transformer (Q-Former)

A lightweight Transformer whose learned queries extract the most text-relevant features from a frozen image encoder.

Drafted by AI · not yet reviewed

How this idea evolved

Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.

  1. 2014

    Let a decoder look back at the most relevant input positions instead of one fixed vector.

    Cites · not yet reviewedcited 5× · §Model Architecture
    “Most competitive neural sequence transduction models have an encoder-decoder structure [5, 2, 35].”
    From Transformer · §Model Architecture
  2. 2017

    Every token attends to every other token, weighted by learned query–key similarity.

    Cites · not yet reviewedcited 3× · §Introduction
    “Transformer, (Vaswani et al. 2017)).”
    From Set Transformer · §Introduction
  3. 2018

    A few learned vectors attend to a large set and summarise it: the idea behind later resamplers and query transformers.

    Cites · not yet reviewedcited 2× · §Methods
    “Attention is a permutation-invariant operation, and this property is preserved by the Perceiver and related models (Lee et al. 2019).”
    From Perceiver · §Methods
  4. 2021

    A small latent array cross-attends to huge inputs, decoupling depth from input size.

    Also draws on: Self-attention (Transformer)

    Cites · not yet reviewedcited 2× · §Approach
    “Similar to Perceiver [48] and DETR [13], we learn a predefined number of latent input queries which are fed to a Transformer and cross-attend to the visual features.”
    From Flamingo · §Approach
  5. 2022

    Learned queries cross-attend to many visual features and compress them to a fixed number of tokens.

    Cites · not yet reviewedcited 5× · §Introduction
    “Frozen (Tsimpoukelli et al. 2021), Flamingo (Alayrac et al. 2022)) resort to an image-to-text generation loss, which we show is insufficient to bridge the modality gap.”
    From BLIP-2 · §Introduction
  6. 2023

    A lightweight Transformer whose learned queries extract the most text-relevant features from a frozen image encoder.

    Also draws on: Multimodal mixture of encoder–decoder (BLIP)

Papers using this

No other paper in this dataset is tagged with it yet.