Paper Lineage
Esc
objectiveIntroduced by CoCa · 2022

Contrastive + captioning in one model

One image–text encoder–decoder trained with both a contrastive and a captioning loss.

Drafted by AI · not yet reviewed

How this idea evolved

Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.

  1. 2018

    Predict future latents and score them against negatives with a contrastive loss.

    Cites · not yet reviewedcited 1× · §Methods
    “This loss takes the same form as the InfoNCE loss Oord et al. 2018, and minimizing it leads to encoders that maximally preserve the mutual information between the true pairs under the representation functions.”
    From ConVIRT · §Methods
  2. 2020

    Pull matching image and text embeddings together and push mismatched pairs apart.

    Also draws on: Joint image–text embedding (Deep Fragment Embeddings)

    No direct link between these papers in this dataset
  3. 2022

    One image–text encoder–decoder trained with both a contrastive and a captioning loss.

    Also draws on: PrefixLM vision-language pre-training (SimVLM)

Papers using this

No other paper in this dataset is tagged with it yet.