architectureIntroduced by Florence · 2021
Vision foundation model
One large image–text model adapted to classification, retrieval, detection, video and more.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that Florence cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Classify into unseen classes by comparing an image with text descriptions of them.
Skip expensive cleaning: a billion noisy alt-text pairs beat curated datasets.
Pull matching image and text embeddings together and push mismatched pairs apart.