Paper Lineage
Esc
architectureIntroduced by BLIP · 2022

Multimodal mixture of encoder–decoder

One model that acts as a contrastive encoder, a matching encoder or a caption decoder.

Drafted by AI · not yet reviewed

How this idea evolved

Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.

  1. 2018

    Predict future latents and score them against negatives with a contrastive loss.

    Cites · not yet reviewedcited 1× · §Methods
    “This loss takes the same form as the InfoNCE loss Oord et al. 2018, and minimizing it leads to encoders that maximally preserve the mutual information between the true pairs under the representation functions.”
    From ConVIRT · §Methods
  2. 2020

    Pull matching image and text embeddings together and push mismatched pairs apart.

    Also draws on: Joint image–text embedding (Deep Fragment Embeddings)

    No direct link between these papers in this dataset
  3. 2021

    Align unimodal embeddings contrastively before fusing them with cross-attention.

    Also draws on: Two-stream co-attention (ViLBERT)

    Cites · not yet reviewedcited 14× · §Related Work
    “Due to the prohibitive expense of acquiring human-annotated texts, most methods (Chen et al. 2020; Li et al. 2020; Li et al. 2021a; Wang et al. 2021; Radford et al. 2021) use image and alt-text pairs crawled from the web (Sharma et al. 2018; Changpinyo et al. 2021; Jia et al. 2021), Despite the use of simple rule-based filters, noise is still prevalent in the web texts.”
    From BLIP · §Related Work
  4. 2022

    One model that acts as a contrastive encoder, a matching encoder or a caption decoder.

    Also draws on: Unified LM via attention masks (UniLM)

Papers using this

No other paper in this dataset is tagged with it yet.