Datasets & data curation
How training and evaluation data were collected, cleaned and scaled.
| Concept | Introduced by | Years | Papers |
|---|---|---|---|
| Dense image–language annotation Human-written regions, attributes and relationships that ground language in images. | Visual Genome | 2016 | 0 |
| Web-scale image–text pairs Hundreds of millions of image–caption pairs scraped from the web. | — | 2016–2022 | 18 |
| Diverse curated text corpora Mix many high-quality text sources to improve LM generalisation. | The Pile | 2020–2022 | 6 |
| CLIP-filtered open datasets Use CLIP similarity to filter web pairs into an open, model-ready dataset. | LAION-400M | 2021 | 0 |