dataIntroduced by ALIGN · 2021
Scaling on raw alt-text
Skip expensive cleaning: a billion noisy alt-text pairs beat curated datasets.
Drafted by AI · not yet reviewed
How this idea evolved
Each step is the idea’s introducing paper. The line between steps is the citation link between those papers.
- 2018
Predict future latents and score them against negatives with a contrastive loss.
Cites · not yet reviewedcited 1× · §Methods“This loss takes the same form as the InfoNCE loss Oord et al. 2018, and minimizing it leads to encoders that maximally preserve the mutual information between the true pairs under the representation functions.”
From ConVIRT · §Methods - 2020
Pull matching image and text embeddings together and push mismatched pairs apart.
Also draws on: Joint image–text embedding (Deep Fragment Embeddings)
No direct link between these papers in this dataset - 2021
Skip expensive cleaning: a billion noisy alt-text pairs beat curated datasets.
Papers using this
- 2014Bahdanau attention
- 2015BookCorpus (books & movies)
- 2019VisualBERT
- 2019VL-BERT
- 2021ALBEF
- 2021SimVLM
- 2021VLMo
- 2021LiT
- 2021Florence
- 2022BLIP
- 2022Flamingo
- 2022CoCa