architectureIntroduced by BLIP · 2022
Multimodal mixture of encoder–decoder
One model that acts as a contrastive encoder, a matching encoder or a caption decoder.
Drafted by AI · not yet reviewed
How this idea evolved
Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.
- 2018
Predict future latents and score them against negatives with a contrastive loss.
Cites · not yet reviewedcited 1× · §Methods“This loss takes the same form as the InfoNCE loss Oord et al. 2018, and minimizing it leads to encoders that maximally preserve the mutual information between the true pairs under the representation functions.”
From ConVIRT · §Methods - 2020
Pull matching image and text embeddings together and push mismatched pairs apart.
Also draws on: Joint image–text embedding (Deep Fragment Embeddings)
No direct link between these papers in this dataset - 2021
Align unimodal embeddings contrastively before fusing them with cross-attention.
Also draws on: Two-stream co-attention (ViLBERT)
Cites · not yet reviewedcited 14× · §Related Work“Due to the prohibitive expense of acquiring human-annotated texts, most methods (Chen et al. 2020; Li et al. 2020; Li et al. 2021a; Wang et al. 2021; Radford et al. 2021) use image and alt-text pairs crawled from the web (Sharma et al. 2018; Changpinyo et al. 2021; Jia et al. 2021), Despite the use of simple rule-based filters, noise is still prevalent in the web texts.”
From BLIP · §Related Work - 2022
One model that acts as a contrastive encoder, a matching encoder or a caption decoder.
Also draws on: Unified LM via attention masks (UniLM)
Papers using this
No other paper in this dataset is tagged with it yet.