architectureIntroduced by ViT-VQGAN · 2021
ViT-based image tokenizer
A ViT-based VQGAN gives better image tokens for autoregressive image modelling.
Drafted by AI · not yet reviewed
How this idea evolved
Suggested from citations. No curator has recorded what this idea builds on yet. These ideas from the same theme come from papers that ViT-VQGAN cites, directly or one step removed. Citation is a fact; the connection between the ideas is not verified.
Model text and image tokens as one stream and sample images from text.
Compress images into a vocabulary of discrete codes so Transformers can model them like text.
Papers using this
- 2022BEiT v2